Direct prediction adds no recovery multiplier.
ECCV 2026
Not All Prediction Targets Keep Training-Free Diffusion Guidance on the Manifold
- 1 Maum AI
- 2 Seoul National University of Science and Technology
* Equal contribution† Corresponding author
In one sentence
Prediction target changes the failure mode.
The gradient-based TFG methods studied here act through an estimate of the clean image. With direct x-prediction, guidance can miss the target while the sample remains realistic. For ε-prediction, clean-image recovery multiplies prediction error by , which diverges as and can send the sample off the manifold.

01 · A practical TFG benchmark
Can guidance hit the target and keep the image realistic?
Training-free guidance steers a pretrained model at inference time without updating its weights. In our benchmark, the model starts from one of 30 ImageNet bird classes and an external guide requests one of 143 species. The benchmark measures whether guidance reaches the requested species and whether the generated distribution remains realistic.
- 1
Parent class
CFG selects one of 30 ImageNet bird classes already represented in the frozen model.
- 2
Species target
The external guide requests one of 143 child species. The model weights do not change.
- 3
Separate evaluation
A second classifier measures validity. Child FID and Precision/Recall measure realism and coverage.
- 4
Guidance sweep
We evaluate the trade-off across guidance strengths rather than report one selected setting.
- Fine-grained targets
- 143 species
- Coarse priors
- 30 parents
- Every operating point
- 9,152 images
- Model coverage
- 6 configs · 3 targets · 2 spaces
Child FID compares all 9,152 guided samples, pooled across 143 targets, with the full fine-grained bird reference set. The pooled score rises when classifier-valid outputs drift from realistic bird imagery.
02 · What failure looks like
A wrong bird can still look like a bird.
Validity only asks whether the evaluation classifier predicts the requested species. Under stronger guidance, a graceful miss remains coherent; a catastrophic miss becomes an artifact that activates the classifier.
1 · most robust
xx-prediction
JiT-H/16 · pixel space
Mostly realistic misses
Strong guidance may miss the requested species while still producing a plausible bird. The image is wrong for the target, but its structure survives.

2 · intermediate
vv-prediction
SiT-XL/2 · latent space
Visible degradation
v-prediction recovers the clean image with a bounded multiplier. Its samples remain recognizable more often than ε-prediction samples, but they degrade more than x-prediction samples.

3 · least robust
εε-prediction
DiT-XL/2 · latent space
Severe off-manifold artifacts
Early recovery errors can grow without bound. Under strong guidance, classifier-friendly textures replace coherent bird anatomy.

These grids show the failure modes of pretrained models. Because the models also differ in architecture, capacity, and operating space, the controlled studies below test the prediction target directly.
03 · Why targets differ
Recovery changes the size of prediction error.
The paper uses the flow-matching convention . Each target can produce a clean-image estimate. ε-prediction is the only one whose recovery formula divides by t.
The high-noise regime
Early denoising steps determine global structure.
TFG needs a clean-image estimate at these steps, even though the current state is still noisy. ε-recovery is singular here, while the manifold-restoring force is weak.
Guidance is most vulnerable here
Its multiplier is at most one.
The multiplier diverges as t → 0.
Recovery multipliers should not be read as equal-error measurements. The base errors δx(t), δv(t), and δε(t) differ. At t=1, z1=x, so v- and ε-recovery return the clean state without using their predicted targets. That endpoint does not rank robustness.


- x-prediction
- 93.3%
- v-prediction
- 21.5%
- ε-prediction
- 0.5%
These are the measured on-manifold rates at D=512. The comparison does not assume equal base errors.
04 · Controlled attribution
Holding the model fixed isolates the target.
The crossed-lines study uses the same residual MLP architecture, data, training budget, guidance, and sampling for all three prediction objectives. The gap grows with ambient dimension.

| Prediction target | Rate | Failure character |
|---|---|---|
| x | 93.3% | Structure preserved |
| v | 21.5% | Bounded but degraded |
| ε | 0.5% | Near-total collapse |
1
Matched networks
Three residual MLPs share the same architecture, data, and training setup across D ∈ {2, 8, 32, 128, 512}. Only the prediction target changes.
2
Latent-space control
DiT and SiT share their architecture family, scale, training data, and operating space. Their ordering is consistent with the target effect.
3
Smaller x-prediction model
JiT-B has 131M diffusion parameters yet reaches C-FID 31.3, ahead of 675M-parameter DiT at 36.7 and SiT at 34.4.
4
Pixel-space control
JiT and PixelFlow both avoid a latent VAE, which provides a complementary comparison within pixel space.
05 · Real-image evidence
Classifier validity can hide manifold damage.
Child FID compares the pooled guided distribution across all 143 targets with the full fine-grained bird reference set. The score worsens when classifier-valid outputs drift from realistic bird imagery.
x · JiT-H
32.9Child FIDv · SiT
34.7Child FIDε · DiT
38.1Child FIDx → ε gap
5.2points at matched validityBenchmark scale: 143 bird species are nested under 30 ImageNet parents. Every plotted point contains 9,152 generated images, with 64 samples per species. The matched operating points are all near 26.6% validity.


| Model | Target | Guidance ρ | Validity | P-FID | C-FID |
|---|---|---|---|---|---|
| JiT-H/16 | x | 3.0 | 26.60% | 6.91 | 32.85 |
| SiT-XL/2 | v | 1.0 | 26.64% | 8.21 | 34.66 |
| DiT-XL/2 | ε | 0.1 | 26.69% | 6.70 | 38.11 |
06 · Broader validation
The pattern repeats, with a smaller gap in style transfer.
We test two more guidance rules, a second fine-grained domain, style transfer, Precision/Recall, and inverse problems.

LGD
JiT-H retains the lowest C-FID frontier.
Paper details ↗
FreeDoM
JiT-H again has the lowest C-FID frontier.
Paper details ↗
Butterfly benchmark
The x, v, ε ordering also appears across 34 butterfly species.
Butterfly data ↗
Style transfer
The separation is smaller than on birds. At ρ=10, JiT-H retains 80% content accuracy while DiT falls to 1.5%.
Paper details ↗
Precision and recall
JiT-H reaches 0.59 recall. DiT peaks at 0.24 precision, stays near 0.49 recall, and loses precision as guidance strengthens.
Paper details ↗| Task | Best LPIPS | ρ |
|---|---|---|
| Gaussian deblur | 0.2140 | 16 |
| 4× super-resolution | 0.1886 | 24 |
Inverse problems
x-prediction reaches the best LPIPS among the evaluated models on both tasks and remains stable over a wider range of ρ. No model beats the degraded-input PSNR baseline, so the result is perceptual rather than a universal reconstruction win.
Reproduction code ↗Resources and citation
Paper, code, data, and citation.
Paper
ECCV 2026 · accepted
ECCV's official listing confirms the title and authors. The manuscript is available on arXiv.
View the ECCV record ↗Open the PDF ↗View the arXiv record ↗Code
Evaluation and inference
The repository contains the crossed-lines study, bird benchmark, and inverse-problem evaluations.
Open the repository ↗Data
Published benchmark resources
Cite
Provisional ECCV BibTeX
@inproceedings{lee2026onmanifold,
author = {Lee, Yunsung and Lee, Hyeongmin},
title = {Not All Prediction Targets Keep Training-Free Diffusion Guidance on the Manifold},
booktitle = {Computer Vision -- ECCV 2026},
year = {2026},
note = {To appear}
}ECCV confirms the title, authors, and acceptance. The paper does not yet have a public Springer volume, page range, or DOI, so the citation omits those fields and uses "To appear."
If clipboard access is unavailable, select and copy the citation directly.