Research
BenchSeg: A Large-Scale Dataset and Benchmark for Multi-View Food Video Segmentation
Overview Research area: Computer vision, specifically food image/video segmentation, dataset and benchmark design, and video object tracking. Technical level: Advanced. The paper assumes familiarity w

- arXiv
- 2601.07581
- Published
- 2026-01-12
- Authors
- Ahmad AlMughrabi, Guillermo Rivo, Carlos Jiménez-Farfán, Umair Haroon, Farid Al-Areqi, Hyunjun Jung, Benjamin Busam, Ricardo Marques, Petia Radeva
AI summary
Overview
Research area: Computer vision, specifically food image/video segmentation, dataset and benchmark design, and video object tracking.
Technical level: Advanced. The paper assumes familiarity with segmentation architectures (FPN, CCNet, SeTR, SAM), video object segmentation with memory modules (XMem2, DEVA), and evaluation metrics such as mAP, IoU, and Recall.
Scope in one sentence: The paper introduces BenchSeg, a benchmark of 25,284 manually annotated frames across 55 dish scenes captured under free 360° camera motion, and uses it to evaluate 20 segmentation models and hybrid segmentation–tracking systems on spatial accuracy, temporal stability, and computational cost.
What This Paper Is About
Food segmentation models trained on static, canonical-view food images (such as FoodSeg103) are used to estimate portion size and nutrients, but they are rarely tested on the handheld, free-motion video clips that real dietary monitoring produces. When such models are applied frame-by-frame to a camera sweeping around a plate, they produce fragmented masks, label flicker, and drift that standard frame-by-frame metrics do not capture.
The goal of this work is to build a dataset and evaluation protocol that exposes those failures: dense per-frame food masks over hemispherical camera trajectories, a suite of 20 baselines trained only on FoodSeg103, and new temporal metrics that quantify stability over time rather than per-frame accuracy alone.
Key Contributions
-
BenchSeg dataset: 25,284 manually annotated frames across 55 dish scenes, aggregated from four public datasets — Nutrition5k (N5k), Vegetables & Fruits (V&F), MetaFood3D (MTF), and FoodKit (FKit) — capturing each dish under free 360° camera motion with hemispherical viewpoint coverage.
-
A unified 20-baseline benchmark: 20 state-of-the-art segmentation models (SAM-based, transformer, CNN, and large multimodal families) are trained solely on FoodSeg103 and evaluated on BenchSeg, both alone and combined with video-memory modules, isolating cross-dataset generalization and robustness to unseen camera poses.
-
Temporal stability metrics: a dedicated evaluation protocol with continuity, flicker rate, IoU drift, and volatility, designed to surface failure modes that remain invisible under standard per-frame evaluation.
-
Annotation quality control and deployment-oriented reporting: masks produced through a polygon labeling interface by 3 annotators with a calibration session, double-annotation of a 4,000-frame subset (approximately 16% of the corpus), and systematic reporting of model size, memory footprint, and inference speed.
Main Findings
-
Static segmenters degrade under novel viewpoints: quantitative and qualitative results show that standard image segmenters degrade sharply when exposed to unfamiliar viewpoints and motion patterns.
-
Memory augmentation preserves temporal consistency: models augmented with video-memory modules maintain temporal consistency across frames, and hybrid 2D-segmenter plus memory-based tracking pipelines, where per-frame masks are propagated through a memory module, achieve the most stable performance across datasets.
-
Best model: the combination SeTR-MLA + XMem2 outperforms prior work, improving over FoodMem by 2.63% mAP.
-
Temporal metrics reveal hidden failure modes: methods with comparable per-frame accuracy can exhibit very different temporal behavior (flickering, discontinuities, gradual degradation), which the authors state is not adequately reflected by conventional tracking or frame-wise IoU-based evaluation.
-
Annotation consistency (mAP ± std, Recall ± std), reported per source dataset over the double-annotated subset: FKit 0.9642 ± 0.0064 and 0.9998 ± 0.0007; MTF 0.9723 ± 0.0055 and 0.9943 ± 0.0010; N5K 0.9270 ± 0.0108 and 0.9997 ± 0.0008; V&F 0.9471 ± 0.0172 and 0.9986 ± 0.0030.
-
Dataset composition: FKit contributes 20,606 images across 21 scenes (981.24 ± 142.69 images per scene, range 715 to 1,209, from chocolate_panettone at 1,209 images to yellow_cane at 715); V&F provides 2,308 images across 11 scenes (209.82 ± 18.76 per scene); MTF provides 1,749 images across 13 scenes (134.54 ± 82.65 per scene, with a bimodal distribution of roughly 200-image and 30-image scenes); N5k provides 621 images across 10 scenes (62.10 ± 1.97 per scene). Per-scene image counts span from 30 frames to over 1,200.
-
Benchmark positioning: the comparison table places BenchSeg (2026) against FoodSeg103 (2021), Nutrition5k (2021), VIPSeg (2022), FoodSAM (2023), V&F (2023), FoodMem (2024), MetaFood3D (2024), MeViS (2025), and FoodKit (2025), distinguishing it by combining ontology, video, protocol, multi-view coverage, 20 baselines, and diagnostics.
-
In-domain FoodSeg103 results (partial): the evaluation table is truncated in the available content at the SeTR row. Reported so far, sorted ascending by mIoU: FPN (ResNet-50) 27.8 mIoU / 38.2 mAcc / 218M; ReLeM-FPN (Transformer) 28.9 / 39.7 / 218M; ReLeM-FPN (LSTM) 29.1 / 39.8 / 218M; CCNet (ResNet-50) 35.5 / 45.3 / 381M; ReLeM-CCNet (Transformer) 36.0 / 46.5 / 381M; ReLeM-CCNet (LSTM) 36.8 / 47.4 / 381M; SeTR (ViT-16/B) 41.3 mIoU and above. The remaining entries, the full BenchSeg partition results (Table 8), and the comparison of model size, runtime, and peak memory for all 20 methods are not present in the available content.
-
Licensing: N5k, V&F, and FKit are CC BY 4.0 (redistribution and commercial use permitted); MTF is CC BY-NC 4.0. For non-commercial sources, only derived annotations and retrieval scripts are distributed.
Methodology in Plain English
The authors define a common three-stage abstraction for video food segmentation. First, keyframe segmentation runs a per-frame segmenter to produce initial masks. Second, temporal propagation spreads those masks to the non-key frames using stored features (a memory module). Third, optional late fusion merges fresh per-frame predictions with propagated masks to correct errors and reduce drift. This lets them compare pure image segmenters directly against hybrids that add a memory tracker under identical evaluation criteria.
The data pipeline re-annotates frames from four existing collections with binary food-versus-background masks. Annotators used a polygon interface (LabelMe), followed written guidelines covering reflections, transparent containers, and overlapping food, and delineated partially obstructed food while excluding utensils, packaging, and tableware. A 4,000-frame subset was labeled independently by two annotators, with a third senior reviewer adjudicating disagreements; agreement was measured with per-image mAP and IoU.
For evaluation, all models — FPN, CCNet, and SeTR among the segmenters, plus ReLeM multimodal pretraining variants, FoodLMM, SegMan, FoodMem, and SAM-style promptable models — were trained exclusively on FoodSeg103 using their official recipes, with no fine-tuning on BenchSeg. Transformer backbones kept ImageNet or ImageNet-21k pretraining. Because FoodSeg103 predicts 103 ingredient classes while BenchSeg uses a single binary foreground mask, the authors map outputs by computing foreground probability as 1 minus the background probability and thresholding at 0.5. For promptable models that emit multiple candidate masks, they select the mask maximizing overlap with the previous frame, resetting every K = 30 frames to limit drift. They evaluate both best-single-mask and union-mask protocols.
Training details: images resized to 2049 × 1024 with a scaling ratio between 0.5 and 2.0 and cropped to 768 × 768; 80k iterations, batch size 8, SGD with momentum 0.9 and weight decay 0.0005, initial learning rate 10⁻³ decayed polynomially with power 0.9. Training used 4 Nvidia H100 GPUs (80 GB VRAM); BenchSeg testing used 1 Nvidia RTX 5090 (32 GB VRAM) and 1 Nvidia RTX 3090 (24 GB VRAM).
Temporal metrics are computed from frame-wise IoU over a sequence. Continuity C_γ is the fraction of consecutive frame pairs whose IoU stays at or above γ = 0.5. Flicker rate FR_δ is the fraction of consecutive pairs where IoU drops by more than δ = 0.2. IoU drift is the mean absolute frame-to-frame IoU difference, and volatility is the standard deviation of IoU across the sequence. All temporal metrics are averaged per scene, then macro-averaged across scenes within each partition.
Why This Matters
Impact on research. The paper argues that benchmarks should diagnose failure, not just rank models. By treating temporal stability as a first-class evaluation dimension alongside spatial accuracy, it provides a protocol that can distinguish methods with identical per-frame scores but different real-world behavior. Releasing 25,284 annotations, train/test splits, and evaluation code for both spatial and temporal metrics gives the field a shared testbed for multi-view food video segmentation.
Real-world applications:
- Automated dietary assessment and calorie/nutrient logging from phone-recorded meal videos.
- Portion-size estimation, where accurate mask boundaries determine volume and therefore nutrient estimates.
- Mobile or embedded food-recognition assistants, where the reported model size, memory, and inference-speed comparisons guide deployment choices.
- Clinical or nutrition-research monitoring of eating behavior, where temporal flicker and mask drift directly affect measurement reliability.
Industry relevance. Food and nutrition apps, smart-kitchen appliances, and health-monitoring platforms all depend on segmentation that survives handheld capture. The paper's explicit reporting of computational efficiency (model size, memory footprint, inference speed) targets deployment-oriented choices rather than leaderboard placement alone. Because BenchSeg aggregates four separately licensed source datasets and releases only derived annotations plus retrieval scripts where redistribution is restricted, it also demonstrates a compliance-aware model for building benchmarks on top of existing corpora.
Future Directions
-
Closing the viewpoint gap at training time. All models were trained only on FoodSeg103 and degrade under free-motion viewpoints; the authors frame this as an isolated generalization test, leaving open whether training on BenchSeg or on multi-view data would remove the degradation.
-
Improving memory propagation under large viewpoint changes. The paper states that open questions remain about how robust memory propagation is when camera pose changes substantially, and about the relative merits of different backbone families.
-
Better temporal metrics and beyond-identity tracking. The authors note that tracking-oriented measures such as identity switches and trajectory continuity target object identity rather than mask-level stability; refining mask-level stability metrics is an explicit direction.
-
Accuracy-stability-cost trade-offs. Systematically characterizing the trade-off between accuracy, temporal stability, and runtime, particularly for mobile and embedded dietary applications, remains an open question the reported efficiency numbers begin to address.
Target Audience
Researchers and graduate students in computer vision working on segmentation, video object segmentation, or benchmark and dataset construction; food-computing and dietary-assessment researchers who need reliable mask quality over video; and applied engineers building mobile or embedded food-recognition and nutrition-estimation products who need both accuracy and efficiency comparisons.
Authors’ abstract
Food image segmentation is a critical task for dietary analysis, enabling accurate estimation of food volume and nutrients. However, current methods suffer from limited multi-view data and poor generalization to new viewpoints. We introduce BenchSeg, a novel multi-view food video segmentation dataset and benchmark. BenchSeg aggregates 55 dish scenes (from Nutrition5k, Vegetables & Fruits, MetaFood3D, and FoodKit) with 25,284 meticulously annotated frames, capturing each dish under free 360° camera motion. We evaluate a diverse set of 20 state-of-the-art segmentation models (e.g., SAM-based, transformer, CNN, and large multimodal) on the existing FoodSeg103 dataset and evaluate them (alone and combined with video-memory modules) on BenchSeg. Quantitative and qualitative results demonstrate that while standard image segmenters degrade sharply under novel viewpoints, memory-augmented methods maintain temporal consistency across frames. Our best model based on a combination of SeTR-MLA+XMem2 outperforms prior work (e.g., improving over FoodMem by ~2.63% mAP), offering new insights into food segmentation and tracking for dietary analysis. In addition to frame-wise spatial accuracy, we introduce a dedicated temporal evaluation protocol that explicitly quantifies segmentation stability over time through continuity, flicker rate, and IoU drift metrics. This allows us to reveal failure modes that remain invisible under standard per-frame evaluations. We release BenchSeg to foster future research. The project page including the dataset annotations and the food segmentation models can be found at https://amughrabi.github.io/benchseg.