Research
Asking like Socrates: Socrates helps VLMs understand remote sensing images
Overview Research area: Remote sensing (RS) vision-language models and multimodal reasoning — specifically, how to make reasoning models genuinely ground their answers in visual evidence from satellit
- arXiv
- 2511.22396
- Published
- 2025-11-27
- Authors
- Run Shao, Ziyu Li, Zhaoyang Zhang, Linrui Xu, Xinran He, Hongyuan Yuan, Bolei He, Yongxing Dai, Yiming Yan, Yijun Chen, Wang Guo, Haifeng Li
AI summary
Overview
Research area: Remote sensing (RS) vision-language models and multimodal reasoning — specifically, how to make reasoning models genuinely ground their answers in visual evidence from satellite and aerial imagery rather than merely narrating a plausible-sounding thought process.
Technical level: Advanced. The paper assumes familiarity with reinforcement learning for language models (GRPO, PPO, reward shaping), supervised fine-tuning, chain-of-thought reasoning, and multi-agent LLM pipelines.
One-sentence scope: The paper diagnoses a "pseudo reasoning" failure mode in remote sensing VLMs, attributes it to a single-pass "Glance Effect," and proposes RS-EoT — an iterative evidence-seeking reasoning paradigm instilled via a Socratic multi-agent data synthesis system (SocraticAgent) plus a two-stage progressive RL pipeline, producing a model called RS-EoT-7B.
Note on completeness: The provided content ends with a supplementary heading, "A empirical find," followed by no text. That section is empty in the supplied material. arXiv identifier: arXiv:2511.22396v2 [cs.CV] 08 Apr 2026; license CC BY 4.0. Code, data, and models are stated to be available at https://geox-lab.github.io/Asking_like_Socrates.
What This Paper Is About
Remote sensing images cover wide spatial extents with large scale variations and sparse, subtle visual clues. The authors observe that reasoning VLMs applied to these images exhibit "pseudo reasoning": they produce explicit reasoning chains, but their performance on RS tasks shows no gain and sometimes degrades below the non-reasoning base model. The paper attributes this to the "Glance Effect" — the model takes a single, coarse perception pass over the image and then reasons from linguistic self-consistency rather than from visual evidence.
The goal is to replace this static, one-glance perception with a language-driven loop in which the model repeatedly revisits the image, asks itself what visual fact it still needs, inspects the image to get that fact, and refines its reasoning accordingly. The paper builds a data synthesis system to create such traces, then uses SFT plus two-stage RL to instill and generalize the behavior in a 7B model.
Key Contributions
-
RS-EoT (Remote Sensing Evidence-of-Thought): A language-driven, iterative visual evidence-seeking reasoning paradigm for remote sensing, in which reasoning is orchestrated by natural language and visual information functions as on-demand evidence within the reasoning chain rather than as a single static global perception.
-
SocraticAgent: A self-play multi-agent system that synthesizes RS-EoT reasoning traces from scratch. It comprises a Reasoner (text-only, instantiated with GPT-5-mini), a Perceiver (image-aware, instantiated with Gemini-2.5-flash), and a Verifier (instantiated with doubao-seed-1.6-thinking). Traces are filtered by requiring the Reasoner — which never sees the image — to reach the correct answer. This produced the RS-EoT-4K dataset, used for SFT cold-starting.
-
A two-stage progressive Grounding+VQA RL strategy: Stage 1 performs RL on fine-grained grounding tasks (DIOR-RSVG and VRSBench) with an IoU-based reward to strengthen evidence-seeking; Stage 2 performs RL on general RS VQA (RS-VQA and FIT-RS) to generalize. To make simple VQA data usable for RL, the authors propose a multiple-choice reconstruction method plus a graded reward function that mitigates reward hacking.
-
RS-EoT-7B: A model trained with the above pipeline that achieves state-of-the-art results reported in the paper across multiple RS VQA and grounding benchmarks, with analyses (case studies, attention dynamics, reward curves, stage-wise ablation) supporting that it mitigates the Glance Effect.
Main Findings
-
Pseudo reasoning is an empirical phenomenon in RS: The paper reports that existing multimodal reasoning models show explicit thinking that can degrade performance below the non-reasoning base model on RS tasks, and that their VQA metrics fluctuate significantly relative to the base model rather than delivering consistent gains.
-
RS-EoT-7B leads on RS VQA benchmarks (Table 1): On RSFG-VQA, RS-EoT-7B scores Avg@5 67.85, Conv@5 68.90, and Pass@5 85.28, versus 62.45 / 62.46 / 77.16 for Qwen2.5-VL. On VRSBench it scores 63.09 / 64.12 / 83.54, and on RSVQA 75.16 / 78.29 / 92.51. On RSFG-SC it reports Scene@acc 64.05 and Object@F1 56.52.
-
RS-EoT-7B leads on grounding (Table 1): On DIOR-RSVG it reports IoU@50 47.00, IoU@70 33.32, and mIoU 45.29; on VRSBench-Ref it reports IoU@50 54.71, IoU@70 32.40, and mIoU 48.04. One number in the table is an exception to the "best across the board" framing: on VRSBench-Ref IoU@70, VHM-RL is reported at 35.40 versus 32.40 for RS-EoT-7B.
-
SocraticAgent traces beat direct distillation on most metrics (Table 3): On FIT-RSFG-VQA with 2K sampled queries, SocraticAgent yields Avg@5 64.82, Conv@5 67.35, Pass@5 85.89, versus Qwen3-VL-plus at 63.78 / 65.15 / 76.87, Doubao-seed-1-6-vision at 51.80 / 17.17 / 66.81, and GLM-4.5V at 66.25 / 66.58 / 77.99. SocraticAgent has the highest Conv@5 and Pass@5; GLM-4.5V has the highest Avg@5.
-
Training stages play complementary roles (Table 4): From the Base model, adding SFT raises VQA (Avg@5 67.20 to 70.73; Pass@5 77.95 to 91.96) but collapses grounding (mIoU 35.64 to 13.95). Adding RL-IoU restores and improves grounding (mIoU 45.57, IoU@50 48.41) while VQA dips slightly (Avg@5 69.51). Adding RL-VQA raises VQA to the best values (Avg@5 75.16, Conv@5 78.29, Pass@5 92.51) while holding grounding near the RL-IoU level (IoU@50 48.23, mIoU 45.52).
-
Attention shows periodic reasoning–perception cycles: Across eight randomly sampled test cases, token-wise attention to image tokens periodically peaks (interpreted as evidence-seeking phases) and falls (interpreted as language reasoning phases).
-
The multiple-choice RL reconstruction stabilizes training: The Stage 2 reward curve rises smoothly from approximately 0.62 to converge around 0.84, which the authors present as evidence that reward hacking is mitigated.
-
Qualitative case studies: In a VQA example the baseline makes one coarse scene interpretation and fails to find an available jet bridge, while RS-EoT-7B first establishes the airport context, searches for an unoccupied gate, and identifies one. In a grounding example the baseline misses a small irregular water body, while RS-EoT-7B performs a global scan, identifies a candidate region in the lower-left, and re-examines the image to rule out other water bodies.
Methodology in Plain English
Step 1 — Define the target behavior. The authors first specify what a good reasoning trace should look like: language plans the steps, and the image is consulted on demand for specific facts, in a repeating loop, until an answer is reached.
Step 2 — Manufacture training traces with two agents talking to each other. Because no existing model naturally produces this pattern, the authors build SocraticAgent. One text-only agent (the Reasoner) never sees the image — it only knows basic metadata such as modality, count, and size — and must reason and ask questions. A second, image-aware agent (the Perceiver) sees the image and answers those questions, but not the original task query. A third model (the Verifier) checks the Reasoner's final answer against ground truth; only correct dialogues are kept as training traces. A six-round dialogue limit is used, producing 4.3K SFT samples.
Step 3 — Force detailed, incremental questioning through self-play prompting. Left alone, the Reasoner tends to forward the whole query to the Perceiver ("jumpy" reasoning), and the Perceiver tends to dump exhaustive, partly irrelevant visual detail. The fix is to tell each agent that its partner is weak: the Reasoner is told the Perceiver cannot understand complex questions, and the Perceiver is told the Reasoner has weak reasoning ability. This pushes the Reasoner toward simple, incremental questions and the Perceiver toward accurate but concise answers.
Step 4 — SFT cold-start. The synthesized RS-EoT-4K traces are used for supervised fine-tuning of Qwen2.5-VL-7B-Instruct using LLaMA-Factory: 5 epochs, learning rate 3×10⁻⁵, AdamW optimizer, cosine learning rate scheduler, global batch size 64, maximum sequence length 4096 tokens, run on 4 A100 GPUs in approximately 40 minutes.
Step 5 — Stage 1 RL on grounding. Using EasyR1 and GRPO, the model is trained on DIOR-RSVG and VRSBench. It must output a bounding box in the format "[x1, y1, x2, y2]"; the reward is the IoU with ground truth, plus a strict format-based reward. Two epochs, batch size 512, learning rate 10⁻⁶.
Step 6 — Stage 2 RL on general VQA, with simple data made harder. Existing RS VQA data is dominated by Yes/No or low-reasoning-intensity questions and is susceptible to reward hacking. The authors exploit the fact that these datasets often map one image to many QA pairs, and rebuild each sample as a multiple-choice question: collect an image and its QA set with 10 < m < 15 pairs, randomly invert the answers of n of them (changing "Yes" to "No," or adding/subtracting a random integer to numeric answers) where 1 < n < m, then ask "Which of the following QA pairs match this remote sensing image?" with the m QA pairs as options. The reward is graded: r_qa = 1 − (1/N) Σ |y_i − ŷ_i|, so choosing a correct option and correctly rejecting an incorrect option both earn credit, while choosing an incorrect option or missing a correct one earns nothing. This symmetrically penalizes misses and false selections and forces the model to verify each option against the image. Also two epochs, batch size 512, learning rate 10⁻⁶, trained on RS-VQA and FIT-RS.
Evaluation setup. Benchmarks are FIT-RSFG-VQA, FIT-RSFG-SC, VRSBench-VQA, and RSVQA for general QA; DIOR-RSVG and VRSBench-Ref for grounding. Metrics are Avg@5, Conv@5, and Pass@5 (average, majority, and any-correct across 5 generated answers), accuracy for classification, F1 for multi-object recognition, and mIoU plus IoU@N for N = 50 and 70. Baselines include WeThink, Vision-R1, R1-OneVision, VLAA-Thinker, VL-Rethinker, MM-Eureka, and the RS-specific Geo-R1 and VHM-RL — all based on Qwen2.5-VL-7B except Geo-R1 and VHM-RL. For the data comparison, distillation is tested from Qwen3-VL-plus, Doubao-seed-1-6-vision, and GLM-4.5V; Gemini and GPT-5 models are excluded because their internal processes are inaccessible.
Why This Matters
Impact on research. The paper reframes a widely reported problem in multimodal reasoning — that reasoning chains sometimes hurt rather than help — as a perception-budget problem rather than a language-model capability problem. Instead of assuming more reasoning tokens automatically help, it argues that reasoning must be able to trigger new perceptual evidence. It also connects reasoning-model training to remote sensing domain constraints (wide spatial extent, large scale variation, sparse and subtle clues) and offers a concrete data-synthesis recipe (a Reasoner that cannot see the image, validated by a Verifier) for creating reasoning traces when no model already exhibits the target behavior. The VQA data reconstruction trick is a general technique for turning simple, reward-hackable datasets into stable RL signal.
Real-world applications (as named or illustrated in the paper):
- Environmental monitoring, using RS imagery interpretation.
- Urban planning, where land-cover and object layout matter.
- Disaster response, where rapid and evidence-grounded scene assessment is needed.
- Grounding tasks such as locating small, irregular features (for example, a water body) and aviation-related inspection such as identifying an available jet bridge or unoccupied gate; the paper also uses ship heading as an illustrative reasoning example.
Industry relevance. The work is a collaboration between Central South University and Baidu Inc. (with Zhejiang University), and the pipeline is deliberately built from off-the-shelf parts — GPT-5-mini, Gemini-2.5-flash, and doubao-seed-1.6-thinking for data synthesis; LLaMA-Factory for SFT; EasyR1 with GRPO for RL; Qwen2.5-VL-7B-Instruct as the base model. The reported SFT stage ran on 4 A100 GPUs in approximately 40 minutes, which is within reach of modest industrial teams. State-of-the-art grounding numbers also matter directly to geospatial production workflows such as object detection and referring expression localization on satellite and aerial imagery.
Future Directions
- Closing the remaining grounding gap: RS-EoT-7B is reported at 32.40 IoU@70 on VRSBench-Ref while VHM-RL is reported at 35.40, so the claim of uniform superiority on grounding does not hold at every threshold metric; understanding when the evidence-seeking loop trades off localization precision is an open question.
- Scaling beyond a 7B backbone: Every model is 7B, based on Qwen2.5-VL-7B-Instruct; whether the RS-EoT pattern persists, strengthens, or becomes redundant at larger scale is not reported.
- Attribution of the source of gains: Table 3 shows SocraticAgent traces leading on Conv@5 (67.35) and Pass@5 (85.89) but trailing GLM-4.5V on Avg@5 (64.82 versus 66.25); isolating which properties of the Socratic traces cause the difference is a natural next study.
- Extending the paradigm beyond remote sensing and beyond the current reward design: The graded multiple-choice reward and the IoU reward are both domain-tailored; whether the same reconstruction strategy transfers to other visual domains with simple QA data, and how far the 6-round SocraticAgent dialogue limit constrains trace quality, are unresolved.
Target Audience
This paper is most useful to researchers and engineers working on multimodal reasoning models, reinforcement learning from verifiable rewards, and vision-language models for Earth observation. It will be particularly valuable to practitioners who need to build domain-specific reasoning models but lack pre-existing reasoning traces for that domain — the SocraticAgent + progressive RL recipe is the
Authors’ abstract
Recent multimodal reasoning models, inspired by DeepSeek-R1, have significantly advanced vision-language systems. However, in remote sensing (RS) tasks, we observe widespread pseudo reasoning: models narrate the process of reasoning rather than genuinely reason toward the correct answer based on visual evidence. We attribute this to the Glance Effect, where a single, coarse perception of large-scale RS imagery results in incomplete understanding and reasoning based on linguistic self-consistency instead of visual evidence. To address this, we propose RS-EoT (Remote Sensing Evidence-of-Thought), a language-driven, iterative visual evidence-seeking paradigm. To instill this paradigm, we propose SocraticAgent, a self-play multi-agent system that synthesizes reasoning traces via alternating cycles of reasoning and visual inspection. To enhance and generalize these patterns, we propose a two-stage progressive RL strategy: first, RL on fine-grained Grounding tasks to enhance RS-EoT capabilities, followed by RL on RS VQA to generalize to broader understanding scenarios. Experiments show RS-EoT achieves state-of-the-art performance on multiple RS VQA and grounding benchmarks. Analyses reveal clear iterative cycles of reasoning and evidence seeking, confirming RS-EoT mitigates the Glance Effect and enables genuine evidence-grounded reasoning. Our code, data, and models are available at https://geox-lab.github.io/Asking_like_Socrates