Skip to content
AI.info

Research

Program-Verified Self-Evolution for Vision-Language Models

Program-Verified Self-Evolution for Vision-Language Models Authors: Ahmed Heakl (LG AI Research, MBZUAI), Sungik Choi (LG AI Research), Moontae Lee (LG AI Research, Australian National University, Uni

Program-Verified Self-Evolution for Vision-Language Models
arXiv
2609.33855
Published
2026-09-27
Authors
Ahmed Heakl, Sungik Choi, Moontae Lee, Salman Khan

AI summary

Program-Verified Self-Evolution for Vision-Language Models

Authors: Ahmed Heakl (LG AI Research, MBZUAI), Sungik Choi (LG AI Research), Moontae Lee (LG AI Research, Australian National University, University of Illinois at Chicago), Salman Khan (MBZUAI) arXiv: 2609.33855v1 [cs.CV], 27 Sep 2026 · License: CC BY 4.0 · Code: https://github.com/ahmedheakl/VQS

Overview

Research area: Self-supervised / self-evolving training of vision-language models (VLMs), specifically the problem of verifying answers when no gold labels exist; also touches on programmatic data generation and reinforcement learning with verifiable rewards.

Technical level: Advanced. The paper assumes familiarity with vision-language architectures, GRPO-style reinforcement learning, schema-constrained decoding, and multimodal benchmark suites, though the core idea (compute answers instead of guessing them) is easy to state.

One-sentence scope: The paper introduces VQS, a framework that replaces majority-vote or model-judge pseudo-labels in self-evolving VLM training with answers computed by fixed programs run over structured image parses, verified claim by claim.

What This Paper Is About

Self-evolving vision-language models improve by generating their own questions from unlabeled images, but because those questions have no gold answers, the model must guess which of its own answers is correct. Prior methods use majority voting over sampled answers or a model judge as a proxy label, and a human evaluation in this paper finds that 24% of majority-vote labels and 18% of model-judge labels were wrong. The goal is to make the training label correct by construction: parse the image into a structured record, let deterministic programs write the question and compute its answer, and use the model only to check the individual visual facts the program read.

Key Contributions

  1. Program-computed QA generation. VQS converts photographs, charts, diagrams, and infographics into structured records (scene graphs, chart tables, diagram node/edge graphs, infographic entries) via schema-constrained parsing, then uses fixed, hand-written templates to write questions and compute their answers rather than estimating answers from model outputs.
  2. Label-free parser improvement. A best-of-K selection algorithm scores K candidate parses by the fraction of their atomic claims the visual checker confirms, and keeps the best parse as an SFT target only if it beats the mean of all K parses by a margin δ and asserts at least m claims. This trains the parser without any labels.
  3. A verified self-evolving training loop. Solver training uses GRPO with a reward of 0.9 × exact match against the computed answer plus 0.1 × format, after three filters: a fact-check, a blind gate (drops questions answerable without the image), and a difficulty band (keeps questions the solver gets right sometimes but not always).
  4. Empirical validation across scales and model families. VQS improves the average over ten benchmarks by up to 3.18 points at the 2B, 4B, and 8B Qwen3-VL scales, beats the strongest baseline at every scale, keeps improving over three training cycles to 3.84 points at 2B, and also transfers to Gemma3-12B-It, InternVL3-8B-Instruct, and Llama-3.2-11B-Vision-Instruct.

Main Findings

  • Superior answer quality. In a human evaluation of 500 shared questions (125 per visual domain), VQS answers were 94.4% correct while retaining 76.0% of questions; majority voting reached 76.4% correctness at 92.0% retention, and a model judge reached 82.2% at 90.8% retention. The conclusion states the error rate drops to 6% (from 24% for majority vote and 18% for the model judge).

  • Largest average gain at every scale. Against each backbone's own base:

Method 2B avg. 4B avg. 8B avg.
Base 59.22 66.70 68.81
VISE (strongest baseline) 60.23 (+1.01) 67.74 (+1.04) 69.47 (+0.66)
VQS (ours) 62.40 (+3.18) 69.01 (+2.32) 71.24 (+2.43)

VQS lifts all ten benchmarks at each scale; VISE lowers LogicVista and MMStar at every scale, while VQS raises both.

  • Biggest gains where the base model is weak. At 2B, VQS gains +7.86 on MMMU (vs +1.75 for VISE) and +7.39 on ScienceQA, benchmarks that require reading a chart, table, or figure and combining several facts. The paper attributes this to the multi-step templates.

  • Pseudo-label accuracy degrades in prior work. VisPlay's estimated pseudo-label accuracy falls from 72% to 61% over three rounds, and R-Zero's falls from 79% to 63%.

  • Per-claim checking beats whole-parse checking for selecting parser targets (150 hand-rated samples):

Parser target Parse Precision Answer Acc. Solver Acc.
Untrained 80.8 86.8 58.0
Random 82.1 91.0 61.5
Whole parse 85.0 94.0 62.0
Per claim 87.6 96.1 62.4

Questions built from an untrained parser even drop the solver below the 2B base, attributed to 13.2% of their answers being wrong.

  • Label-free training genuinely improves the parser. Measured as F1 against human annotations, it gains 5.0 points on GQA, 6.6 on ChartQA, and 7.4 on AI2D-RST. Diagrams are the weakest domain before training (75.2) and remain hardest after training (82.6).

  • Ablations rank the components. Removing parser training is the largest hit (parser precision 87.6 → 80.8, solver accuracy 62.4 → 58.0). Removing the fact-checker is second (solver 62.4 → 60.9 with parser precision unchanged). Removing the blind gate (61.8), the difficulty band (62.1), or constrained decoding (61.9, parser precision 85.3) each costs 0.3 to 0.6 points.

  • Multi-read template families drive the gains. The largest per-family gains are chart differences (+7.4) and diagram connectivity (+5.8); the smallest are object existence (−0.2) and attribute lookup (+0.4), likely because the base model already handles single reads.

  • Gains scale with template coverage. Average gain rises from 1.9 points with 10 template families to 3.18 with all 70, and is still rising; reasoning benchmarks gain 4.7 points against 2.4 for perception benchmarks at every family count.

  • Gains shrink with model size, tracking how much the model changes. The share of benchmark answers changed by RL falls from 7.82% at 2B to 5.58% at 4B and 4.93% at 8B. All three scales reach a training reward of about 0.80, so larger models learn the generated task but transfer less of it.

  • More cycles keep helping. Average gain over base grows from 3.18 points after cycle 1 to 3.41 and 3.84 after cycles 2 and 3; the third cycle adds more than the second (0.43 vs 0.23 points). After three cycles, knowledge and reasoning benchmarks gain 5.94 points, led by ScienceQA (+8.53) and MMMU (+8.44), while perception gains 2.62 and general 2.24. Nine of ten benchmarks improve every cycle; MMStar peaks after cycle 1 but still ends 0.93 points above base.

  • Question design choices matter. Middle-difficulty families work best (62.4) over hard (61.1) and easy (60.8); fixed templates beat model-written programs (62.4 vs 61.4, because only 85.3% of generated programs matched their question); ordering questions within each family by hop count beats shuffled order (62.4 vs 61.3); and reselecting questions during training gives no gain (62.3).

  • Paraphrasing rarely breaks the question. Across 437 questions from 70 template families, the rewritten question still matched its program in 99.8% of cases at 2B, 100% at 4B, and 99.1% at 8B. All four 8B errors swap the two years in a yes/no comparison, making the stored answer wrong.

  • The checker is not merely rubber-stamping the parser. On 1,417 atomic claims extracted from 150 parses (1,170 parser claims, 247 deliberately corrupted traps), the checker rejects 244 of the 247 traps. Agreement with other checkers tracks model size more than family: Qwen3-VL-8B κ = 0.837, InternVL3.5-8B κ = 0.810, Qwen2.5-VL-7B κ = 0.807, SmolVLM2-2.2B κ = 0.660. Among accepted claims, at least two other checkers reject 26.7% of diagram arrows, 8.3% of chart values, and none of the object-existence claims.

  • Transfer beyond Qwen. VQS improves all ten benchmarks for every tested family, averaging +1.77 on Gemma3-12B-It, +1.51 on InternVL3-8B-Instruct, and +1.84 on Llama-3.2-11B-Vision-Instruct.

Methodology in Plain English

The only training input is a set of unlabeled images (12,000 sourced from Venkatraman et al. (2026), with their original questions, answers, and labels discarded). A single pretrained model plays three roles: parser, solver, and checker.

Step 1 — Parse the image. The parser converts each image into a JSON record whose shape depends on the domain: photos become objects with names, attributes, boxes, and relations; charts become series of (category, printed value, numeric value) points plus axis ticks; diagrams become nodes and directed arrows; infographics become entries with text, number, and unit. A constrained decoder enforces the schema so output is always well formed.

Step 2 — Write questions and compute answers with programs. A fixed library of 70 hand-written templates traverses the parse and returns a question, its answer, and a hop count (the number of operations performed). Templates combine operations like filtering a set, reading a field, counting, taking a max or min, following a relation, ranking, comparing, doing arithmetic, reading a trend, counting a node's arrows, and converting a box to a coarse position. The parser then paraphrases the question into natural wording while the computed answer stays fixed.

Step 3 — Filter the questions. A question survives only if three conditions hold: a fact-check confirms every atomic claim the template read (checked one short claim at a time, with a yes/no answer over four new tokens); a blind gate shows that a text-only version of the model, answering J = 4 times without the image, does not get more than λ = 0.5 of them right; and a difficulty band shows the solver gets the question right at least once but not every time across 8 rollouts at temperature 1.0 with a 256-token budget.

Step 4 — Train the solver. Using GRPO with a LoRA adapter of rank 64 (vision modules excluded, so no visual parameters are trained), the reward is 0.9 × exact match with the computed answer plus 0.1 × format compliance. Rollouts are scored against the computed answer, so the correct answer is reinforced even when the model's own majority opinion was wrong.

Step 5 — Train the parser without labels. For each image, the parser samples K = 4 parses at temperature 1.0 and top-p 0.95 (up to 2048 tokens). Each parse is scored by the fraction of its claims the checker confirms. The best parse is kept as an SFT target only if it beats the mean of all K parses by δ = 0.10 and asserts at least m = 4 claims. This turned 12,000 images into 4,792 targets. The parser is then fully fine-tuned with the vision tower frozen (2 epochs, 284 steps, learning rate 1e-5).

Hardware was 8×AMD MI210 cards; GRPO used EasyR1 and parser fine-tuning used LLaMA-Factory. Evaluation used lmms-eval on ten benchmarks — GQA, OK-VQA, InfoVQA, ScienceQA, MMMU, MMBench, EmbSpatial, LogicVista, MMStar, and SEEDBench — with default hyperparameters for each. In later cycles, the parser and solver are retrained on the same images with new template-generated QA pairs, and questions are reweighted by a Gaussian weight w(s) = exp(−((s−0.5)/0.15)²) so that questions solved about half the time appear most often.

Why This Matters

Impact on research. The paper attacks the weakest link in self-evolving multimodal training: the pseudo-label. It shows that consistency-based proxies (majority vote, model judges, transformation agreement) accept many wrong answers and get worse as the loop runs, and offers a different design — make the answer computable rather than estimated. The claim-level verification mechanism is also reusable as a data-quality tool independent of the training loop, and the finding that middle-difficulty, multi-hop, image-dependent questions carry the most learning signal is directly actionable for anyone building synthetic multimodal data.

Real-world applications (grounded in the paper's four supported domains):

  • Chart and infographic question answering (InfoVQA-style): automatically reading series values, axis ticks, and printed numbers from business dashboards, financial reports, or scientific figures, where answers must be exact rather than plausible.
  • Technical and scientific diagram understanding: following directed arrows and node connectivity in flowcharts, circuit diagrams, and process diagrams, which the paper identifies as the hardest domain and the one where the checker most needs improvement.
  • Educational tutoring and assessment: generating and grading questions over figures and tables, matching the ScienceQA and MMMU benchmark categories where VQS gains the most (+8.53 and +8.44 after three cycles).
  • Accessible image description and visual grounding: producing verified object, attribute, and spatial-relation records for photos, which underpin scene-graph-based perception tasks such as GQA and EmbSpatial.

Industry relevance. The method needs no human annotation at any stage and uses unlabeled images only, which reduces the cost of domain adaptation for chart-heavy, diagram-heavy, or infographic-heavy verticals. It is demonstrated at the 2B, 4B, and 8B scales — model sizes practical to serve — and the best results come at 2B, where the absolute gain is largest and the share of changed benchmark answers is highest (7.82%). The multi-cycle recipe, which keeps improving to 3.84 points, offers a way to keep extracting value from the same unlabeled image pool.

Future Directions

  • Strengthen the checker on the hardest claim types. At least two other checkers reject 26.7% of accepted diagram arrows and 8.3% of accepted chart values, and diagrams remain the weakest parser domain after training (82.6 F1). The paper notes that an independent check requires a larger model rather than merely a different family.
  • Close the transfer gap at larger model scales. Training reward is roughly 0.80 at 2B, 4B, and 8B, but benchmark gains fall from 3.18 to 2.43 and the share of changed answers falls from 7.82% to 4.93%. Why larger models transfer less is left open.
  • Extend template coverage and domain schemas. Average gain was still rising at all 70 template families (3.18 points, against 1.9 with 10 families), suggesting more families, and possibly more image domains beyond photo, chart, diagram, and infographic, could add further gains.
  • Improve question generation beyond fixed templates. Model-written programs scored 61.4 against 62.4 for fixed templates, and only 85.3% of generated questions matched their underlying program, so better generators remain an open target if they can raise that matching rate.

Target Audience

Researchers and engineers working on multimodal large language models, synthetic data generation, and reinforcement learning with verifiable rewards — especially those building self-improving or self-play training loops without human labels. It will also interest practitioners in document intelligence, chart and diagram understanding, and educational AI who need answer accuracy that is auditable rather than merely self-consistent

Authors’ abstract

Self-evolving vision-language models train on questions they generate from unlabeled images. Since these questions have no gold answers, prior methods label them by majority vote over sampled answers or by a model judge. In a human evaluation, we find that 24\% of majority-vote labels and 18\% of model-judge labels produced during self-evolution are wrong. To address this problem, we present Verifiable QA Generation for Self-Evolving Models (VQS), which changes how the model judges answers. Instead of voting on an answer, the model parses each image into a structured record, such as a scene graph, a chart table, or a diagram graph. Fixed programs then write a question from the record and compute its answer. The model still acts as a visual checker, but it only confirms the individual facts the program reads, one short claim at a time. These claim-level checks select the parser's training targets, so the parser also improves without labels. Human raters find 94\% of VQS answers correct, against 76\% for majority voting. Across ten benchmarks, VQS improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales and outperforms the strongest self-evolving baseline at each. Gains keep growing over three training rounds, reaching 3.84 points at 2B. Code is released at https://github.com/ahmedheakl/VQS

Read the original paper