Journal Article10.48550/arxiv.2408.05710
Efficient Diffusion Transformer with Step-wise Dynamic Attention Mediators
Yifan Pu,Zhuofan Xia,Jiayi Guo,Dongchen Han,Qixiu Li,Duo Li,Yuhui Yuan,Jihong Li,Yizeng Han,Shiji Song,Gao Huang,Xiu Li +11 more
- 11 Aug 2024
TL;DR: This paper proposes an efficient diffusion transformer framework with step-wise dynamic attention mediators, reducing query-key interaction redundancy and computational complexity, achieving state-of-the-art image quality with a FID score of 2.01 and lower inference cost.
read more
Abstract: This paper identifies significant redundancy in the query-key interactions within self-attention mechanisms of diffusion transformer models, particularly during the early stages of denoising diffusion steps. In response to this observation, we present a novel diffusion transformer framework incorporating an additional set of mediator tokens to engage with queries and keys separately. By modulating the number of mediator tokens during the denoising generation phases, our model initiates the denoising process with a precise, non-ambiguous stage and gradually transitions to a phase enriched with detail. Concurrently, integrating mediator tokens simplifies the attention module's complexity to a linear scale, enhancing the efficiency of global attention processes. Additionally, we propose a time-step dynamic mediator token adjustment mechanism that further decreases the required computational FLOPs for generation, simultaneously facilitating the generation of high-quality images within the constraints of varied inference budgets. Extensive experiments demonstrate that the proposed method can improve the generated image quality while also reducing the inference cost of diffusion transformers. When integrated with the recent work SiT, our method achieves a state-of-the-art FID score of 2.01. The source code is available at https://github.com/LeapLabTHU/Attention-Mediators.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Figures

Fig. 1: (a) shows the JSD-based redundancy score defined in Sec. 3.2 evaluated on DiTS/2 model along with diffusion time steps. The score is computed over 32 samples and averaged by different attention heads in every layer. (b) shows the same redundancy score of all the 12 layers of SiT-S/2 model with the SDE sampler. 
Fig. 2: Ablation for optimized mediator token adjustment schedule. (a) Trade-off between FID-50K and FLOPs. (b) Trade-off between sFID-50K and FLOPs. ![Fig. 3: Main Results of the proposed method in 256×256 resolution. Each string of red dots is obtained by adjusting the mediator token number with optimized thresholds. (a) Comparison with DiT [58] and SiT [52]; (b) Zoomed in results around SiT-B/2.](/figures/figure3-1-m9r17bvs2p1h.png)
Fig. 3: Main Results of the proposed method in 256×256 resolution. Each string of red dots is obtained by adjusting the mediator token number with optimized thresholds. (a) Comparison with DiT [58] and SiT [52]; (b) Zoomed in results around SiT-B/2. 
Table 1: Effectiveness of static mediator tokens. n is the mediator tokens number. 
Fig. 5: Sampled images by SiT-XL/2 models endowed with our method trained on ImageNet 256×256 resolution with cfg=4.0. 
Table 3: Effectiveness of static mediator tokens. n is the mediator tokens number.
Citations
Dynamic Tuning Towards Parameter and Inference Efficiency for ViT Adaptation
Wangbo Zhao,Jiasheng Tang,Yizeng Han,Yibing Song,Kai Wang,Gao Huang,Fan Wang,Yang You +7 more
TL;DR: Dynamic Tuning (DyT) improves both parameter and inference efficiency for ViT adaptation by dynamically skipping redundant computations and introducing an enhanced adapter.
References
•Proceedings Article
Adam: A Method for Stochastic Optimization
Diederik P. Kingma,Jimmy Ba +1 more
- 01 Jan 2015
TL;DR: This work introduces Adam, an algorithm for first-order gradient-based optimization of stochastic objective functions, based on adaptive estimates of lower-order moments, and provides a regret bound on the convergence rate that is comparable to the best known results under the online convex optimization framework.
138.5K
•Posted Content
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
TL;DR: A new language representation model, BERT, designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers, which can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of tasks.
81.7K
•Posted Content
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy,Lucas Beyer,Alexander Kolesnikov,Dirk Weissenborn,Xiaohua Zhai,Thomas Unterthiner,Mostafa Dehghani,Matthias Minderer,Georg Heigold,Sylvain Gelly,Jakob Uszkoreit,Neil Houlsby +11 more
TL;DR: Vision Transformer (ViT) attains excellent results compared to state-of-the-art convolutional networks while requiring substantially fewer computational resources to train.
•Proceedings Article
Language Models are Few-Shot Learners
Tom B. Brown,Benjamin Mann,Nick Ryder,Melanie Subbiah,Jared Kaplan,Prafulla Dhariwal,Arvind Neelakantan,Pranav Shyam,Girish Sastry,Amanda Askell,Sandhini Agarwal,Ariel Herbert-Voss,Gretchen Krueger,Thomas Henighan,Rewon Child,Aditya Ramesh,Daniel M. Ziegler,Jeffrey Wu,Clemens Winter,Christopher Hesse,Mark Chen,Eric Sigler,Mateusz Litwin,Scott Gray,Benjamin Chess,Jack Clark,Christopher Berner,Samuel McCandlish,Alec Radford,Ilya Sutskever,Dario Amodei +30 more
- 28 May 2020
TL;DR: GPT-3 achieves strong performance on many NLP datasets, including translation, question-answering, and cloze tasks, as well as several tasks that require on-the-fly reasoning or domain adaptation, such as unscrambling words, using a novel word in a sentence, or performing 3-digit arithmetic.
•Posted Content
Decoupled Weight Decay Regularization
Ilya Loshchilov,Frank Hutter +1 more
TL;DR: This work proposes a simple modification to recover the original formulation of weight decay regularization by decoupling the weight decay from the optimization steps taken w.r.t. the loss function, and provides empirical evidence that this modification substantially improves Adam's generalization performance.
14.4K