Attention Alignment Regularization [3] is Essential for Texture Accuracy
LLeffa: Aligns intermediate layer keys & queries to map worn regions directly to canonical positions, preventing logo distortion and texture warping.
We investigate and measure the impact on the model performance along three key design axes:
Our final Dual-UNet architecture.
LLeffa: Aligns intermediate layer keys & queries to map worn regions directly to canonical positions, preventing logo distortion and texture warping.
Given a large architecture, an end-to-end joint training of both UNets causes training instability and metric degradation. We design a curriculum to progressively enhance both networks.
Backbone selection dictates CreationNet's generative capacity, while using a base model over inpainting proves crucial. Masking directly impacts low-level fine-grained features, but has negligible effect on global guidance.
for VTOFF because SSIM/LPIPS heavily penalize harmless shape length variations under occlusion while ignoring fine texture loss.
Quantitatively evaluated on VITON-HD and DressCode datasets, our framework achieves state-of-the-art performance with a drop of 9.5% on the primary metric DISTS and competitive performance on LPIPS, FID, KID, and SSIM, providing both stronger baselines and insights to guide future Virtual Try-Off research.
| Method | Resolution | SSIM ↑ | LPIPS ↓ | DISTS ↓ | FID ↓ | KID ↓ |
|---|---|---|---|---|---|---|
| TryOffDiff [1] (HD) | 512 × 384 | 75.02 | 28.52 | 22.32 | 24.56 | 9.52 |
| Try-Off-Anyone [2] (HD) | 512 × 384 | 72.35 | 34.08 | 22.11 | 11.57 | 2.01 |
| Ours (HD) | 512 × 384 | 74.70 | 28.75 | 20.10 | 9.73 | 1.82 |
| IGR [3] (HD) | 1024 × 768 | 78.95 | 29.46 | 20.45 | 13.14 | 2.97 |
| Ours (HD) | 1024 × 768 | 76.04 | 31.41 | 19.59 | 10.69 | 2.31 |
| TryOffDiff [1] (DC)☆ | 512 × 384 | 80.8 | 31.6 | 21.6 | 17.1 | 4.7 |
| Ours (DC)☆ | 512 × 384 | 75.53 | 32.61 | 20.85 | 12.30 | 2.12 |
☆ DISTS is computed at 341 × 256 resolution following the TryOffDiff evaluation protocol.
@inproceedings{truong2026matters,
title={What Matters in Virtual Try-Off? Dual-UNet Diffusion Model For Garment Reconstruction},
author={Truong, Loc-Phat and Madadi, Meysam and Escalera, Sergio},
year={2026},
eprint={2604.08716},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2604.08716}
}