ECCV 2026

Not All Prediction Targets Keep Training-Free Diffusion Guidance on the Manifold

Yunsung Lee¹*Hyeongmin Lee²*†

  1. 1 Maum AI
  2. 2 Seoul National University of Science and Technology

* Equal contribution† Corresponding author

sung@maum.aihyeongmin.lee@seoultech.ac.kr

In one sentence

Prediction target changes the failure mode.

The gradient-based TFG methods studied here act through an estimate of the clean image. With direct x-prediction, guidance can miss the target while the sample remains realistic. For ε-prediction, clean-image recovery multiplies prediction error by (1−t)/t, which diverges as t→0 and can send the sample off the manifold.

Diagram and examples showing x-prediction moving along the data manifold, epsilon-prediction collapsing off it, and v-prediction failing more gracefully
A guided sample may hit the target, miss while remaining realistic, or leave the data manifold entirely. Prediction target changes which failure occurs.

01 · A practical TFG benchmark

Can guidance hit the target and keep the image realistic?

Training-free guidance steers a pretrained model at inference time without updating its weights. In our benchmark, the model starts from one of 30 ImageNet bird classes and an external guide requests one of 143 species. The benchmark measures whether guidance reaches the requested species and whether the generated distribution remains realistic.

  1. 1

    Parent class

    CFG selects one of 30 ImageNet bird classes already represented in the frozen model.

  2. 2

    Species target

    The external guide requests one of 143 child species. The model weights do not change.

  3. 3

    Separate evaluation

    A second classifier measures validity. Child FID and Precision/Recall measure realism and coverage.

  4. 4

    Guidance sweep

    We evaluate the trade-off across guidance strengths rather than report one selected setting.

Fine-grained targets
143 species
Coarse priors
30 parents
Every operating point
9,152 images
Model coverage
6 configs · 3 targets · 2 spaces

Child FID compares all 9,152 guided samples, pooled across 143 targets, with the full fine-grained bird reference set. The pooled score rises when classifier-valid outputs drift from realistic bird imagery.

02 · What failure looks like

A wrong bird can still look like a bird.

Validity only asks whether the evaluation classifier predicts the requested species. Under stronger guidance, a graceful miss remains coherent; a catastrophic miss becomes an artifact that activates the classifier.

  1. 1 · most robust

    xx-prediction

    JiT-H/16 · pixel space

    Mostly realistic misses

    Strong guidance may miss the requested species while still producing a plausible bird. The image is wrong for the target, but its structure survives.

    Grid of JiT x-prediction bird samples under strong guidance; most remain coherent natural birds even when some miss the requested species
  2. 2 · intermediate

    vv-prediction

    SiT-XL/2 · latent space

    Visible degradation

    v-prediction recovers the clean image with a bounded multiplier. Its samples remain recognizable more often than ε-prediction samples, but they degrade more than x-prediction samples.

    Grid of SiT v-prediction bird samples under strong guidance; structure is less stable than x-prediction but more coherent than epsilon-prediction
  3. 3 · least robust

    εε-prediction

    DiT-XL/2 · latent space

    Severe off-manifold artifacts

    Early recovery errors can grow without bound. Under strong guidance, classifier-friendly textures replace coherent bird anatomy.

    Grid of DiT epsilon-prediction bird samples under strong guidance; many contain severe artifacts and collapsed bird-like textures

These grids show the failure modes of pretrained models. Because the models also differ in architecture, capacity, and operating space, the controlled studies below test the prediction target directly.

03 · Why targets differ

Recovery changes the size of prediction error.

The paper uses the flow-matching convention zt=tx+(1−t)ε. Each target can produce a clean-image estimate. ε-prediction is the only one whose recovery formula divides by t.

The high-noise regime

Early denoising steps determine global structure.

TFG needs a clean-image estimate at these steps, even though the current state is still noisy. ε-recovery is singular here, while the manifold-restoring force is weak.

xDirect
‖x^(x)−x‖=δx

Direct prediction adds no recovery multiplier.

vBounded
‖x^(v)−x‖=(1−t)δv

Its multiplier is at most one.

εSingular
‖x^(ε)−x‖=1−ttδε

The multiplier diverges as t → 0.

Recovery multipliers should not be read as equal-error measurements. The base errors δx(t), δv(t), and δε(t) differ. At t=1, z1=x, so v- and ε-recovery return the clean state without using their predicted targets. That endpoint does not rank robustness.

Paper Figure 1(a): x-prediction follows the data manifold from a noisy state toward the target region, while v- and epsilon-prediction can depart from the manifold
Guidance acts through the recovered clean image. Direct x-prediction follows the manifold toward the target. An unstable recovered estimate can send the trajectory away from it.
On-manifold rate versus ambient dimension; x-prediction stays high while v-prediction declines and epsilon-prediction approaches zero
The networks, data, training, guidance, and sampling are identical. Only the prediction objective changes.
x-prediction
93.3%
v-prediction
21.5%
ε-prediction
0.5%

These are the measured on-manifold rates at D=512. The comparison does not assume equal base errors.

04 · Controlled attribution

Holding the model fixed isolates the target.

The crossed-lines study uses the same residual MLP architecture, data, training budget, guidance, and sampling for all three prediction objectives. The gap grows with ambient dimension.

Five rows of crossed-lines results from dimension 2 through 512; x-prediction stays concentrated on the target line while v-prediction spreads and epsilon-prediction largely leaves the line at high dimension
Crossed-lines under DPS guidance with strength 10 and 100 Euler steps. Each target uses the same 256-hidden, five-block residual MLP design. At D=512, the on-manifold rates are 93.3%, 21.5%, and 0.5% for x-, v-, and ε-prediction.
On-manifold rate at D = 512
Prediction targetRateFailure character
x93.3%Structure preserved
v21.5%Bounded but degraded
ε0.5%Near-total collapse

1

Matched networks

Three residual MLPs share the same architecture, data, and training setup across D ∈ {2, 8, 32, 128, 512}. Only the prediction target changes.

2

Latent-space control

DiT and SiT share their architecture family, scale, training data, and operating space. Their ordering is consistent with the target effect.

3

Smaller x-prediction model

JiT-B has 131M diffusion parameters yet reaches C-FID 31.3, ahead of 675M-parameter DiT at 36.7 and SiT at 34.4.

4

Pixel-space control

JiT and PixelFlow both avoid a latent VAE, which provides a complementary comparison within pixel space.

05 · Real-image evidence

Classifier validity can hide manifold damage.

Child FID compares the pooled guided distribution across all 143 targets with the full fine-grained bird reference set. The score worsens when classifier-valid outputs drift from realistic bird imagery.

x · JiT-H

32.9Child FID

v · SiT

34.7Child FID

ε · DiT

38.1Child FID

x → ε gap

5.2points at matched validity

Benchmark scale: 143 bird species are nested under 30 ImageNet parents. Every plotted point contains 9,152 generated images, with 64 samples per species. The matched operating points are all near 26.6% validity.

Plot of parent FID against classifier validity; several models reach similar validity while their parent FID trade-offs differ
Validity records whether the evaluation classifier predicts the requested species. Similar validity does not imply similar visual fidelity.
Plot of parent FID against Child FID; JiT x-prediction traces the lowest Child FID frontier, SiT v-prediction is intermediate, and DiT epsilon-prediction remains higher
Child FID reveals the quality gap at matched validity: 32.85 for JiT-H, 34.66 for SiT, and 38.11 for DiT. Lower is better.
Matched-validity operating points reported in the paper
ModelTargetGuidance ρValidityP-FIDC-FID
JiT-H/16x3.026.60%6.9132.85
SiT-XL/2v1.026.64%8.2134.66
DiT-XL/2ε0.126.69%6.7038.11

06 · Broader validation

The pattern repeats, with a smaller gap in style transfer.

We test two more guidance rules, a second fine-grained domain, style transfer, Precision/Recall, and inverse problems.

LGD bird benchmark frontier with JiT-H x-prediction attaining the lowest Child FID
LGD on the fine-grained bird benchmark.

LGD

JiT-H retains the lowest C-FID frontier.

Paper details ↗
FreeDoM bird benchmark frontier with JiT-H x-prediction attaining the lowest Child FID
FreeDoM on the fine-grained bird benchmark.

FreeDoM

JiT-H again has the lowest C-FID frontier.

Paper details ↗
Butterfly benchmark P-FID versus Child FID plot, with JiT-H x-prediction on the lowest frontier
Thirty-four butterfly species under six ImageNet parents.

Butterfly benchmark

The x, v, ε ordering also appears across 34 butterfly species.

Butterfly data ↗
Style-transfer plot of Gram loss against content validity; x-prediction retains content accuracy more strongly at high guidance
Four WikiArt styles by 100 ImageNet classes, 400 images per setting.

Style transfer

The separation is smaller than on birds. At ρ=10, JiT-H retains 80% content accuracy while DiT falls to 1.5%.

Paper details ↗
Precision versus recall plot; JiT-H reaches recall 0.59 while DiT remains near 0.49 and loses precision under stronger guidance
PRDC separates fidelity from coverage and exposes mode-collapse behavior.

Precision and recall

JiT-H reaches 0.59 recall. DiT peaks at 0.24 precision, stays near 0.49 recall, and loses precision as guidance strengthens.

Paper details ↗
Best x-prediction perceptual results
TaskBest LPIPSρ
Gaussian deblur0.214016
4× super-resolution0.188624

Inverse problems

x-prediction reaches the best LPIPS among the evaluated models on both tasks and remains stable over a wider range of ρ. No model beats the degraded-input PSNR baseline, so the result is perceptual rather than a universal reconstruction win.

Reproduction code ↗

Resources and citation

Paper, code, data, and citation.

Code

Evaluation and inference

The repository contains the crossed-lines study, bird benchmark, and inverse-problem evaluations.

Open the repository ↗

Cite

Provisional ECCV BibTeX

@inproceedings{lee2026onmanifold,
  author    = {Lee, Yunsung and Lee, Hyeongmin},
  title     = {Not All Prediction Targets Keep Training-Free Diffusion Guidance on the Manifold},
  booktitle = {Computer Vision -- ECCV 2026},
  year      = {2026},
  note      = {To appear}
}

ECCV confirms the title, authors, and acceptance. The paper does not yet have a public Springer volume, page range, or DOI, so the citation omits those fields and uses "To appear."

If clipboard access is unavailable, select and copy the citation directly.