Skip to content
AI.info

Research

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection Overview Research area: Computer vision / vision-language models — specifically the intersection of dense image captioning and pixel-

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection
arXiv
2609.19143
Published
2026-09-16
Authors
Sara Pieri, Evangelos Kazakos, Shizhe Chen, Josef Sivic, Cordelia Schmid

AI summary

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection

Overview

  • Research area: Computer vision / vision-language models — specifically the intersection of dense image captioning and pixel-level segmentation (panoptic grounded captioning).
  • Technical level: Advanced. The high-level idea is intuitive, but the paper assumes familiarity with VLMs, DETR-style set prediction, Hungarian matching, panoptic segmentation metrics, and LoRA finetuning.
  • Scope: Introduces PanoCaps, a human-annotated benchmark with near-complete pixel coverage, and PANORAMA, a model that grounds caption phrases by selecting from a phrase-conditioned pool of mask proposals rather than decoding masks directly.

What This Paper Is About

Modern vision-language models can write fluent, detailed image captions, but they struggle to reliably point at the exact pixels their words refer to. The "panoptic grounded captioning" task asks a model to describe an entire scene — foreground objects and background stuff alike — while attaching a pixel-accurate mask to every referring phrase. The authors argue that progress is blocked by two things: existing datasets trade off annotation coverage against quality, and asking an autoregressive language model to predict masks directly overloads it with a spatial task it was never designed for. They address both.

Key Contributions

  1. PanoCaps benchmark. A human-annotated panoptic grounded captioning dataset of 3,470 images (2,070 train / 420 val / 980 test), built from COCONut, ADE20K, and VIPSeg. It contains ~34.1K masks, 31.3K grounded entities, ~9 entities per image, 17.9K unique noun phrases, and roughly 99% pixel coverage. Captions are free-form scene descriptions with each phrase explicitly linked to one or more masks.

  2. An evaluation protocol and the gPQ metric. A phrase-mask matching procedure that handles free-form text via exact, lemma-based, WordNet-synonym, and sentence-embedding matching, followed by Hungarian assignment. The generalized Panoptic Quality (gPQ) metric extends standard PQ to free-form phrases by weighting each matched pair by both its mask IoU and its textual similarity, and penalizing unmatched predictions (FP) and unmatched ground-truth regions (FN).

  3. The PANORAMA model. A grounding formulation in which the VLM emits a [SEG] token per phrase, a "concept bridge" projects the token's hidden state into a 256-dim concept vector, a pretrained SAM 3 proposal model is conditioned on that vector to produce a pool of 200 candidate masks, and a learned match scorer selects the proposals corresponding to the phrase. Grounding is thereby reduced to selection rather than mask generation.

  4. Strong and transferable results. Best overall grounding on PanoCaps, competitive or superior to specialized models on grounded conversation generation, referring expression segmentation, and generalized referring segmentation, plus the finding that finetuning existing baselines on PanoCaps consistently improves their grounding.

Main Findings

  • PanoCaps closes the density/quality gap. Prior datasets are either sparse (GranD-f grounds only a handful of entities, leaving large parts of the scene described but ungrounded) or dense but noisy (COCONut-PanCap produces generic statements, incorrect phrase-mask links, and masks limited to COCO categories). PanoCaps provides both density and verified precision.

  • Human annotation includes natural multi-referencing. 9.8% of phrases refer to several masks (e.g., grouping instances of a category), and 12.5% of masks are referenced by more than one phrase — a structure that single-mask-per-phrase models cannot represent.

  • Decoupling referent identification from boundary prediction works. Because the VLM decides what is being referred to and the pretrained segmenter decides where its edges are, PANORAMA inherits the mask quality of a large-scale segmenter instead of learning mask decoding from scratch. The paper reports this beats [SEG]-decoding approaches that force one representation to encode both the referent and its spatial extent.

  • Set-valued output handles singular, plural, and absent referents. A single [SEG] token per phrase can yield zero, one, or many selected masks, since each proposal is thresholded independently at θ = 0.5. This contrasts with methods that collapse multiple instances into one binary mask or require a separate token per instance.

  • Conditioning the proposal pool on the phrase matters. The ablation compares PANORAMA against a GROUNDHOG-style approach that selects from an image-level, phrase-independent proposal pool, and against using SAM 3's native CLIP text encoder in place of the concept vector — the latter would condition the segmenter on phrase text alone rather than the VLM's contextualized representation, and would require a 354M-parameter text tower per phrase.

  • Finetuning on PanoCaps benefits other models. Baselines finetuned on PanoCaps for 10 epochs show consistent grounding improvements, indicating the benchmark's supervision is genuinely useful, not merely diagnostic.

  • Training scale is modest. The full 957K-sample mixture trains in ~5 hours (2B) and ~7 hours (4B) on 16 NVIDIA H100 GPUs, which makes the approach comparatively accessible.

  • Run-to-run variability is small. Training the 2B model with three seeds yields standard deviations of 0.6 gPQ, 1.0 AP50, and 0.1 mean RefCOCO cIoU — so differences of that magnitude should be read as ties.

Methodology in Plain English

The core insight is a division of labor. Instead of asking a language model to draw masks — a job it is bad at — the authors ask it only to identify what is being talked about.

The pipeline runs in four stages. First, a Qwen3-VL model writes a caption where every referring phrase is wrapped in delimiters and followed by a special [SEG] token. Second, the hidden state of each [SEG] token is squeezed through a small learned projection into a 256-dimensional "concept vector." Third, that concept vector is fed into a pretrained SAM 3 segmenter, which uses it to generate 200 candidate masks specifically conditioned on that concept. Fourth, a lightweight scorer compares each candidate against the concept vector and keeps the ones scoring above 0.5.

The key design choice is that SAM 3's own text encoder is removed entirely. Normally SAM 3 segments instances of a noun phrase by encoding that phrase with a CLIP-style text tower. Here the concept vector replaces it, which means the segmenter sees the VLM's full contextual understanding — the image plus the preceding sentence — rather than the phrase in isolation.

Training uses four losses: standard next-token cross-entropy for the caption, a sigmoid focal loss on the selection scores, mask losses (BCE plus Dice) on matched proposals, and a semantic loss applied to the segmenter's pretrained semantic head (discarded at inference). Positive and negative proposal labels come from Hungarian matching, where the cost combines mask overlap (Dice) with the match score itself, so proposals the scorer already likes are preferred when overlaps are similar. Only the LoRA adapters, token embeddings, concept bridge, fusion encoder, and match scorer are trained — both vision encoders stay frozen.

The training mixture spans 957K samples across five capability groups: dense grounded captioning (PanoCaps plus a regenerated COCONut-PanCap), grounded conversation generation, referring segmentation, generalized referring segmentation, and multi-granularity segmentation. The COCONut-PanCap portion required significant cleanup, since its captions are noisy and its phrase-mask correspondences incomplete.

Why This Matters

The paper's argument is that spatial grounding and descriptive richness should not be a trade-off, and it backs that up with both a benchmark and a method. This matters because many downstream systems need to act on what a model says, not just read it — an inaccurate mask can cause a robot to grab the wrong object or an assistive system to describe the wrong thing.

Research impact:

  • Provides a benchmark where caption completeness and grounding precision are measured jointly rather than separately, which existing datasets could not support.
  • The gPQ metric gives the community a graded, free-form-compatible extension of the widely used Panoptic Quality, replacing binary match counting.
  • The phrase-conditioned proposal selection formulation is a reusable architectural pattern that any VLM-plus-segmenter pairing could adopt, and it sidesteps the need for mask tokenizers or task-specific pixel decoders.

Real-world applications:

  • Robotic manipulation and embodied agents — robots that follow natural-language instructions need to know exactly which object and which pixels an instruction refers to before grasping or moving.
  • Assistive technology for blind and low-vision users — scene description systems must accurately localize what they mention, and multi-region grounding lets them say "the two chairs on your left" and mean it.
  • Image editing and content creation — text-driven editing ("remove the chili powder") depends on precise, entity-level mask proposals that follow the user's phrasing.
  • Autonomous driving and scene monitoring — full-scene description with per-region grounding supports both foreground actors and background context, which is how driving scenes are actually annotated.

Industry relevance: The method reuses pretrained components rather than training segmenters from scratch, trains in hours on commodity H100 nodes, and drops in as a finetuning stage on an existing VLM family. That profile fits organizations that already have a VLM in production and want to add grounding without building a segmentation stack.

Future Directions

  • Scaling the annotation. PanoCaps costs roughly 1,530 human hours for 3,470 images (median ~15 min labeling plus ~12 min review each). Extending to larger and more domain-diverse corpora will require either much cheaper verification or a semi-automatic pipeline that preserves the current quality bar.

  • Reducing dependence on the pretrained segmenter. PANORAMA's mask quality is bounded by SAM 3's proposal pool. It is an open question how the approach performs when the base segmenter struggles — small objects, unusual domains, or overlapping instances where 200 queries may not cover the relevant region.

  • Improving the free-form matching procedure. Textual agreement currently relies on a cascade of exact, lemma, WordNet, and embedding matches. This heuristic chain is the weakest link in both training and evaluation for phrases that are paraphrases or refer to compound concepts.

  • Unifying more of the grounding landscape. The paper positions referring segmentation, generalized referring segmentation, and category-level grounding as separate points in a two-axis space (provided vs. generated phrases; single vs. multi-region targets). Whether a single model can cover that whole space without task-specific tuning is left open.

  • Cross-lingual and cross-domain generalization. The paper shows qualitative results on out-of-domain scenes but does not systematically measure how far the grounding capability transfers beyond the source segmentation corpora.

Target Audience

Researchers and engineers working on vision-language models, grounded captioning, referring segmentation, or panoptic segmentation will find the benchmark and the proposal-selection architecture directly relevant. It is also useful for practitioners building embodied agents, assistive vision tools, or text-driven image editing systems who need pixel-accurate, entity-level localization from natural language. Readers should come prepared with background in segmentation metrics and multimodal model training; the paper's framing is clear enough that the core idea survives without deep familiarity with the internals, but the evaluation protocol and ablation design reward a technical reader.

Authors’ abstract

Intelligent systems that act in the world require image understanding that is both comprehensive and spatially grounded. Current vision-language models (VLMs) can generate fluent and detailed image captions, but reliably associating them with image pixels remains challenging. Existing methods that combine dense captioning with pixel-level grounding often produce either incomplete descriptions or inaccurate segmentation masks. We study this problem through panoptic grounded captioning, a task that requires a VLM to describe both foreground objects and background regions while grounding each referring phrase with pixel-level masks. We make three contributions. First, we introduce PanoCaps, a human-annotated benchmark constructed from panoptic segmentation datasets. It provides dense captions with near-complete pixel coverage and image-text alignments at the entity level, supporting both training and evaluation. We further propose a phrase-mask matching protocol and a generalized Panoptic Quality (gPQ) metric that jointly evaluates textual and mask agreement. Second, we formulate phrase grounding as selection from a phrase-conditioned pool of mask proposals and introduce PANORAMA, a VLM that conditions a pretrained segmenter on contextualized phrase representations to obtain candidate masks and learns to select those corresponding to each phrase. Training this interface jointly with caption generation enables PANORAMA to produce high-quality masks while allowing each phrase to refer to a single region or multiple instances. Third, PANORAMA achieves the best overall grounding on PanoCaps and matches or exceeds specialized models across several pixel-level grounding tasks. Experiments show that our method produces precise entity-level segmentations while maintaining detailed, mask-consistent captions. Code, data and models are available at https://www.di.ens.fr/willow/research/panorama/.

Read the original paper