Skip to content
AI.info

Research

A Multimodal, Multitask System for Generating E Commerce Text Listings from Images

Overview Research area: Multimodal machine learning — vision-language models applied to e-commerce product content generation. Technical level: Advanced. The work assumes familiarity with vision-langu

A Multimodal, Multitask System for Generating E Commerce Text Listings from Images
arXiv
2510.21835
Published
2025-10-22
Authors
Nayan Kumar Singh

AI summary

Overview

Research area: Multimodal machine learning — vision-language models applied to e-commerce product content generation. Technical level: Advanced. The work assumes familiarity with vision-language models, multi-task learning, model fine-tuning, and text generation metrics. Scope: The paper proposes and evaluates a single-image, end-to-end system that jointly predicts product attributes and price and then generates a grounded text listing, measuring both factual accuracy and generation latency against single-task and non-hierarchical baselines.

What This Paper Is About

Retailers currently rely on people to write product names and descriptions, which is slow and costly, and this paper examines whether generative AI can automate it. Existing vision-to-language models tend to invent facts, and training separate models for each product attribute ignores the relationships between those attributes. The goal is a single system that turns one product image into a text listing whose claims are factually grounded in what the image actually shows.

Key Contributions

  1. Multi-task fine-tuning of a vision encoder. A single vision backbone is trained jointly on attribute prediction tasks (color, hemline, neck style) and on price regression, rather than using siloed single-task models.
  2. A hierarchical generation process. The model's own predicted attributes are embedded into a prompt that is fed to the text decoder, so generation is conditioned on explicit predicted facts rather than on the image alone.
  3. An end-to-end multimodal, multitask system. The work integrates attribute prediction, price regression, and listing generation into one pipeline that takes a single image as input.
  4. An empirical comparison against single-task and ablated alternatives. The design choices are evaluated against independent price regression, independent attribute classification, a non-hierarchical ablation, and a comparable direct vision-to-language model.

Main Findings

  • Multi-task learning improves regression: The multitask approach outperforms independent price regression with a 3.6% better R2 value.
  • Multi-task learning improves classification: The same approach improves attribute classification by 6.6% in F1 score over the single-task alternative.
  • Hierarchical generation sharply reduces hallucination: The hallucination rate drops from 12.7% to 7.1%, a 44.5% relative reduction, compared with a non-hierarchical ablation.
  • Hierarchical generation is faster: Latency of the autoregressive text generation process falls by a factor of 3.5 relative to a direct vision-to-language model of similar size.
  • A stated trade-off in text quality: The model scores 3.5% worse than the direct vision-to-language model on ROUGE-L, which the author identifies as a minor caveat.

Methodology in Plain English

The system takes a product image and produces a text listing. Two design decisions distinguish it from a straightforward image-to-text model.

First, instead of training one model to guess the color, another to guess the hemline, another to guess the neck style, and yet another to estimate price, the researchers train one shared vision backbone on all of these targets at the same time. The reasoning is that these properties are related to one another, so learning them jointly should produce better predictions than learning them in isolation.

Second, the system generates text in stages. Rather than asking a text decoder to look at the image and describe it, the system first produces its attribute predictions, then inserts those predictions into a prompt that is passed to the decoder. The decoder therefore writes from an explicit list of predicted facts rather than from raw visual features, which is the mechanism the paper credits for lowering hallucination.

The abstract reports comparisons against single-task models, against a version of the system without the hierarchical step, and against a direct vision-to-language model. It does not describe the dataset, the training setup, or the architecture specifics, so those details are not available here.

Why This Matters

Research impact. The paper argues that factual grounding can be improved through pipeline and training design — multi-task learning plus attribute-conditioned prompting — rather than only through scaling a model. It also suggests that joint training on related prediction targets is a useful way to make one backbone serve several downstream tasks, which is relevant to anyone building multimodal systems where separate models have become unwieldy.

Real-world applications.

  • Automating product name and description writing for retail catalogs, reducing the manual authoring burden the paper identifies.
  • Generating attributes and prices from images for sellers onboarding new inventory with minimal metadata.
  • Producing consistent, standardized listings across a large catalog where human-written copy varies in quality and format.
  • Reducing the review effort needed to check AI-generated copy, since grounded listings contain fewer invented claims.

Industry relevance. Hallucination is a practical blocker for deploying generative AI in commerce, where an incorrect claim about a product carries legal, financial, and reputational risk. A 44.5% relative reduction in hallucination, combined with a 3.5x latency improvement, speaks directly to the cost and reliability concerns that determine whether such a system can be deployed at catalog scale. The multitask results also address the operational inefficiency of maintaining many separate single-task models.

Future Directions

  • Generalization beyond the tested attributes. The named attributes (color, hemline, neck style) are apparel-specific. Whether the hierarchical approach transfers to other product categories, or to attribute sets that are not enumerated in advance, is left open.
  • Closing the text-quality gap. The 3.5% ROUGE-L deficit against a direct vision-to-language model raises the question of whether grounding and fluency can both be maximized, or whether the trade-off is inherent.
  • Broader evaluation of factual grounding. The abstract reports hallucination rate and ROUGE-L but does not describe the evaluation protocol. How hallucination is defined and measured, and whether human raters were involved, would clarify how much weight the headline number should carry.
  • Extending the multitask targets. The abstract mentions two task types, attribute classification and price regression. Whether additional objectives can be added to the same backbone without degrading the reported gains is an obvious next question.

Target Audience

Researchers working on multimodal generation, vision-language model grounding, and multi-task learning will find the architectural proposals and ablation design most directly useful. Practitioners building e-commerce catalog or listing-automation systems will benefit from the hallucination and latency results, which speak to deployment constraints. Readers looking for a template of how to combine multi-task fine-tuning with attribute-conditioned prompting will find the two-proposal structure the clearest takeaway.

Authors’ abstract

Manually generating catchy descriptions and names is labor intensive and a slow process for retailers. Although generative AI provides an automation solution in form of Vision to Language Models (VLM), the current VLMs are prone to factual "hallucinations". Siloed, single task models are not only inefficient but also fail to capture interdependent relationships between features. To address these challenges, we propose an end to end, multi task system that generates factually grounded textual listings from a single image. The contributions of this study are two proposals for the model architecture. First, application of multi task learning approach for fine tuning a vision encoder where a single vision backbone is jointly trained on attribute prediction such as color, hemline and neck style and price regression. Second, introduction of a hierarchical generation process where the model's own predicted attributes are embedded in a prompt and fed to the text decoder to improve factual consistency. The experiments demonstrate the superiority of this architecture. The multi tasking approach outperforms both the independent price regression, with a 3.6% better R2 Value and attribute classification, with a 6.6% improvement F1 score. Critically, the hierarchical generation process proves highly effective, slashing the factual hallucination rate from 12.7% to 7.1%, a 44.5% relative reduction, compared to a non hierarchical ablation. The hierarchical approach also reduces the latency of the autoregressive text generation process by a factor of 3.5 when compared to direct vision to language model of similar size. One minor caveat is that the model does perform 3.5% worse than direct vision-to-language model on ROUGE-L score.

Read the original paper