Research
Diffusion Is Your Friend in Show, Suggest and Tell
Diffusion Is Your Friend in Show, Suggest and Tell Overview Research area: Computer Vision, specifically image captioning, bridging discrete denoising diffusion models and autoregressive (AR) language

- arXiv
- 2512.10038
- Published
- 2025-12-10
- Authors
- Jia Cheng Hu, Roberto Cavicchioli, Alessandro Capotondi
AI summary
Diffusion Is Your Friend in Show, Suggest and TellOverview
- Research area: Computer Vision, specifically image captioning, bridging discrete denoising diffusion models and autoregressive (AR) language generation.
- Technical level: Intermediate. The paper assumes familiarity with transformer encoder-decoder architectures and the basics of diffusion models, but its central idea is explained through plain architecture diagrams and analogies.
- Scope (1 sentence): The paper proposes using a discrete diffusion model as an auxiliary "suggestion" provider that feeds token hints to a standard autoregressive captioner, and evaluates the resulting Show, Suggest and Tell (SST) architecture on the MS-COCO 2014 dataset.
What This Paper Is About
Diffusion models have produced strong results in generative computer vision, but in text-related tasks such as image captioning they still fail to beat standard autoregressive models and at best match them. Instead of trying to replace autoregressive captioners, this paper asks whether a diffusion model can serve as a helper that suggests likely words to an autoregressive model, combining diffusion's bidirectional refinement with the strong linguistic structure of causal generation. The authors build and test SST to see whether this hybrid role produces measurable gains in caption quality.
Key Contributions
- A new paradigm for diffusion models: The paper proposes adopting diffusion models as Suggestion Modules that support autoregressive generation rather than competing with it, a role the authors describe as currently underexplored.
- The SST architecture: Based on that suggestion module, the authors present Show, Suggest and Tell (SST), a hybrid non-autoregressive/autoregressive captioning model that reports State-of-the-Art results on the MS-COCO 2014 test set among similar models trained with Cross-Entropy loss.
- Extensive validation experiments: The paper runs ablation studies covering suggestion-module configuration, integration strategy, diffusion framework choice, and single-word versus n-gram suggestions, and shows the approach generalizes across captioning models (ExpansionNet and a standard Transformer) and diffusion strategies.
- An oracle study of suggestion potential: By replacing the suggestion module with an Oracle that samples from the ground-truth token set with a preservation probability, the authors quantify the upper bound of benefit that a stronger suggestion module could deliver.
Main Findings
- State-of-the-Art CIDEr-D on COCO: SST achieves 125.1 CIDEr-D on the MS-COCO 2014 test set without Reinforcement Learning, outperforming the reported autoregressive State-of-the-Art by 1.5 points and the reported diffusion State-of-the-Art by 2.5 points. Its full reported scores are Bleu1 78.3, Bleu2 62.6, Bleu3 48.8, Bleu4 37.6, METEOR 29.2, ROUGE 58.3, SPICE 22.8.
- Multiple references beat a single reference for suggestions: Training the suggestion module on all five unique-token references per image consistently raises recall compared to training on one caption per image, which the authors attribute to the network "forgetting" when trained on a single reference.
- Small hidden size is sufficient: A suggestion module with hidden size H = 128 performs comparably to H = 512; the authors kept H = 128 for subsequent configurations.
- Restricting to nouns and verbs hurts: Limiting suggestions to nouns and verbs (POS tags NN, NNS, VBP, VBZ, CD, VBG via Stanford's POS Tagger) caused a significant drop in recall, suggesting the most semantically meaningful words are also the hardest to predict.
- Diffusion beats direct prediction for suggestions: A diffusion-free "Direct Prediction" baseline reached 83.0% precision but only 26.3% recall (F1 39.2%), producing mostly safe words, while Reparametrized Diffusion reached 62.0% precision, 45.4% recall (F1 51.8%). Absorbing Diffusion reached 62.7% / 44.9% (F1 51.6%). Analog-Bit Diffusion without a length penalty reached 23.0% precision / 82.6% recall (F1 34.3%), improving to 47.6% / 53.7% (F1 49.9%) with a length penalty.
- Limited suggestion access is better than full access: Feeding all suggested tokens to the captioner via cross-attention in decoder layers (case A) or encoder layers (case B) reduced CIDEr-D to 122.2 and 122.8 respectively, versus 124.9 for the adopted Reduce + Concatenation into decoder layers strategy (case D). Feeding the pooled suggestion to encoder layers (case C) also proved ineffective, yielding 122.9.
- F1 correlates with caption quality: The authors report a positive correlation between the suggestion module's F1 score and description quality (summarized mainly by CIDEr-D), with the approach improving a Transformer baseline from 117.0 to 120.6 CIDEr-D and ExpansionNet v2 from 123.3 to 124.9 CIDEr-D.
- N-grams do not help: Training the suggestion module on n-grams of size 2 or 3 produced similar captioning results to single tokens; for example, N-gram=3 with 5 references per image yielded 124.2 CIDEr-D versus 125.1 for N-gram=1.
- Oracle study shows headroom: With an Oracle suggestion provider, quality rises monotonically with the preservation probability ρ: BLEU4 and CIDEr-D reach a maximum of 40.1 and 132.6 at ρ = 1. Even at ρ = 0.25 the Oracle reaches 126.6 CIDEr-D, outperforming SST, which the authors attribute to the Oracle's ability to occasionally supply correct tokens where the suggestion module struggles due to epistemic uncertainty.
Methodology in Plain English
The authors split the captioning job into two cooperating parts. The first part, the suggestion module, is a discrete denoising diffusion model built from a standard transformer encoder-decoder sitting on top of a pre-trained, frozen Swin-Transformer Large visual backbone. It learns to generate the set of tokens it expects to appear in a human description for an image, starting from noise and progressively denoising toward a set of words. Since the order of tokens in a set does not matter, this module can process suggestions bidirectionally. An additional length decoder predicts how many suggestion tokens to produce, and it is trained jointly with the diffusion objective through a length loss.
The second part is an ordinary autoregressive captioner, for which the authors use the encoder-decoder refinement layers of ExpansionNet. To pass suggestions across, they deliberately avoid giving the captioner the full set of suggested tokens. Instead, a "Reduce" layer computes a weighted average of all suggested token embeddings into one single vector, which is concatenated onto the decoder's input at each refinement step. The stated rationale is that a weak suggestion module giving only a few correct but incomplete tokens would otherwise act as noise or push the model toward a copy-paste behavior and higher exposure bias; a compressed representation forces the captioner to use suggestions as a guide rather than a source to copy from.
Training and evaluation used the MS-COCO dataset with Karpathy's split: 113,000 training image-description pairs, 5,000 for validation, and an equal number for testing. The suggestion model was trained on 565k human-annotated descriptions (five per image), with captions lowercased, punctuation removed, and words occurring fewer than 5 times filtered out (vocabulary size 10,000). The autoregressive model uses three encoder-decoder layers with hidden size 512, and the suggestion module uses three layers with hidden size 128. The authors adopted the learning rate and optimization settings of ExpansionNet v2 but trained for 20 epochs, and provided suggestions to the captioner only 50% of the time during training (always at inference). At inference, the autoregressive model uses beam search with beam size 3, and the suggestion module performs a maximum of 20 diffusion steps. The autoregressive model is trained with standard Cross-Entropy loss. Evaluation uses BLEU, ROUGE, CIDEr, METEOR and SPICE for captions, and precision and recall for suggestion proficiency.
Why This Matters
The paper argues that diffusion models in sequence tasks have been framed narrowly as cheaper inference alternatives to autoregressive models, an advantage that is eroded by faster autoregressive research. Reframing diffusion as an auxiliary suggestion provider sidesteps the accuracy gap entirely and points to a direction the authors describe as underexplored and promising, particularly with respect to explainable AI.
Real-world applications (grounded in what the paper discusses or implies):
- Assistive and descriptive image captioning: The suggestion mechanism lets the model propose more descriptive attributes, such as adjectives the authors note do not appear in the final prediction, which can support richer automatic descriptions.
- Uncertainty monitoring in deployed systems: The authors propose that a well-performing suggestion module could be used to monitor epistemic uncertainty, which they describe as particularly useful in real-world applications.
- Explainable AI: Because suggestions are visible intermediate outputs, they provide a form of transparency into why a caption was generated, a point the authors raise in their conclusion.
- Foundation for multimodal LLM extensions: The authors position this work as groundwork that can be extended to Multimodal LLMs and massive datasets, addressing what they describe as the lack of diffusion models in the LLM scene.
Industry relevance: The work targets architecture principles rather than scale, deliberately comparing models of standard size (around 100K parameters, as stated in the paper) trained on medium-sized datasets, and the authors cite environmental concerns as a reason for this choice. The code is stated to be available at https://github.com/jchenghu/show_suggest_tell. The work was funded by the European Union's Horizon 2020 programme dAIEDGE (G.A. No 101120726) and supported by the CINECA Italian consortium under project ISCRA-ShareGPT.
Future Directions
- Improving the suggestion module's F1 on semantically meaningful words: The authors calculate that if the suggestion module could achieve a higher F1 score on nouns and verbs than 50%, the captioner could benefit enormously — potentially allowing direct copying from suggestions and even higher scores.
- Monitoring epistemic uncertainty: The authors propose exploring the suggestion module as a transparency tool that flags what a model does and does not know.
- Addressing suggestion imbalance and missed obvious words: The paper notes that some images receive more suggestions than others, hinting at dataset or linguistic bias, and that in some cases even obvious tokens such as "women" are missing — an issue attributed to the infancy of the research topic and to the fact that discrete diffusion frameworks were originally designed for sequences while the suggestion problem is framed as set prediction.
- Scaling to Multimodal LLMs and larger datasets: The authors state the principle should be extendable to LLMs and massive corpora, providing a path for diffusion models to enter the LLM scene.
Target Audience
Researchers and graduate students working on image captioning, discrete diffusion models, or non-autoregressive sequence generation, particularly those interested in hybrid architectures. It is also relevant to practitioners who want a concrete, relatively small-scale example of combining diffusion with autoregressive models, and to readers interested in explainability and uncertainty monitoring in vision-language systems. Readers looking for state-of-the-art results achieved through reinforcement learning, massive training corpora, or Vision Language Models composed of billions of parameters will not find them here, since the paper explicitly scopes itself to standard-sized models trained on medium-sized datasets.
Authors’ abstract
Diffusion Denoising models demonstrated impressive results across generative Computer Vision tasks, but they still fail to outperform standard autoregressive solutions in the discrete domain, and only match them at best. In this work, we propose a different paradigm by adopting diffusion models to provide suggestions to the autoregressive generation rather than replacing them. By doing so, we combine the bidirectional and refining capabilities of the former with the strong linguistic structure provided by the latter. To showcase its effectiveness, we present Show, Suggest and Tell (SST), which achieves State-of-the-Art results on COCO, among models in a similar setting. In particular, SST achieves 125.1 CIDEr-D on the COCO dataset without Reinforcement Learning, outperforming both autoregressive and diffusion model State-of-the-Art results by 1.5 and 2.5 points. On top of the strong results, we performed extensive experiments to validate the proposal and analyze the impact of the suggestion module. Results demonstrate a positive correlation between suggestion and caption quality, overall indicating a currently underexplored but promising research direction. Code will be available at: https://github.com/jchenghu/show\_suggest\_tell.