Research
Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?
Overview Research area: Computer vision and multimodal large language models (MLLMs), specifically chain-of-thought (CoT) reasoning and reinforcement fine-tuning (RFT) applied to fine-grained visual c

- arXiv
- 2601.06993
- Published
- 2026-01-11
- Authors
- Jie Zhu, Yiyang Su, Xiaoming Liu
AI summary
Overview
Research area: Computer vision and multimodal large language models (MLLMs), specifically chain-of-thought (CoT) reasoning and reinforcement fine-tuning (RFT) applied to fine-grained visual classification (FGVC).
Technical level: Advanced. The paper assumes familiarity with GRPO optimization, reward shaping, LoRA, supervised fine-tuning, and open-ended visual question answering evaluation.
Scope: A systematic investigation of when and why textual reasoning helps or hurts MLLMs on fine-grained visual perception, together with a reinforcement fine-tuning framework (ReFine-RFT) that constrains reasoning length while optimizing accuracy-oriented rewards.
What This Paper Is About
Multi-modal large language models are strong generalists but still struggle to tell apart subordinate-level categories that differ only in subtle visual detail, such as car models, plant species, or pet breeds. Chain-of-thought reasoning is widely credited with improving performance on math and coding, but prior work has reported that adding explicit textual reasoning can actually lower accuracy on visual perception tasks. This paper re-examines that issue far more broadly, across zero-shot evaluation and several training paradigms, and asks whether the problem is reasoning itself or the way current methods use it.
Key Contributions
- An empirical characterization of the "Cost of Thinking" in FGVC: verbose CoT systematically degrades MLLM performance, and reasoning length is identified as the key factor, based on zero-shot evaluation, RFT dynamics, thinking-length-controlled RFT, and SFT with GPT-4o-generated CoT data.
- MRN (Multi-reward Normalization), a plug-and-play normalization method that independently normalizes each heterogeneous reward signal before aggregation, rather than normalizing a summed scalar reward as in standard GRPO.
- ReFine-RFT, a reinforcement fine-tuning framework that combines an ensemble, semantically-aware reward with MRN to constrain reasoning length while supplying dense accuracy-oriented feedback.
- State-of-the-art results across multiple FGVC benchmarks with a 2B backbone trained on 4-shot data, including gains over Visual-RFT and over Finedefics-8B trained on the full FGVC datasets.
Main Findings
-
CoT degrades zero-shot FGVC accuracy: Averaged over four benchmarks, Qwen2-VL-2B falls from 60.5 (answer-only) to 55.9 (CoT), Qwen2-VL-7B from 61.2 to 57.8, Qwen2.5-VL-7B from 57.7 to 51.8, InternVL2.5-8B from 28.9 to 26.8, and InternVL3-8B from 29.3 to 25.9. The paper describes this as an average drop of 3–6% for non-reasoning models. Individual model–dataset pairs do not all follow the average: for example, Qwen2-VL-2B on Pets-37 rises from 56.4 (answer-only) to 66.4 (CoT). The reasoning model R1-OneVision-7B-RL scores 54.8 average and generates CoT-style output even under the answer-only prompt, so its answer-only numbers are omitted.
-
Reasoning collapse during RFT: Tracking completion length across RFT on all four FGVC datasets shows average reasoning length trending steadily downward, stabilizing at a range shorter than the zero-shot length of the base model, while accuracy improves. The paper calls this implicit suppression of verbose reasoning "reasoning collapse," and notes it may also result from reward hacking because no explicit length constraint was imposed.
-
Longer reasoning lowers accuracy under explicit control: When thinking length is gradually limited from [0, 20] to [60, 80] during RFT, classification accuracy declines across all FGVC datasets. Shorter reasoning traces yield higher accuracy.
-
Answer-only beats CoT in supervised fine-tuning: Using GPT-4o-generated high-quality CoT data, SFT-AO with LoRA reaches an 80.2 average versus 72.0 for SFT-CoT with LoRA, indicating the degradation is not merely a quality problem with the reasoning traces.
-
Fine-tuning comparisons: LoRA outperforms fully fine-tuning under both SFT and RFT. Visual-RFT with LoRA averages 82.9. ReFine-RFT-AO averages 85.2 (78.7 Aircrafts-102, 81.4 Flowers-102, 93.1 Cars-196, 87.6 Pets-37) and ReFine-RFT-CoT averages 86.5 (79.3, 81.0, 97.1, 88.6), reported as relative improvements of +3.7%, +6.9%, +1.4% and +2.6% over Visual-RFT with LoRA. No-Thinking-RFT (fully fine-tuned) reports 71.2 on Flowers-102 and 86.1 on Pets-37.
-
Prompt style matters less than reasoning length: No-Thinking-RFT uses an answer-only-style prompt and Visual-RFT uses a CoT-style prompt, yet they perform similarly; within ReFine-RFT, the CoT variant is only slightly better than the answer-only variant. The authors conclude that reasoning-length control, not prompt style, is the primary determinant.
-
MRN improves accuracy as a plug-and-play module: Adding MRN to GRPO yields gains of +1.1%/+2.1%/+0.4% (Aircrafts-102, Flowers-102, Cars-196) under full fine-tuning and +0.7%/+1.5%/+0.6% under LoRA. MRN also produces consistently higher reward values and notably lower reward standard deviation than GRPO on Aircrafts-102.
-
Larger LoRA capacity helps in few-shot FGVC: With r=16, α=32 the model scores 71.3/64.6/94.0; with r=32, α=64 it scores 75.4/70.4/95.0; with r=64, α=128 it scores 76.3/75.6/96.3, surpassing the fully fine-tuned baseline.
-
Ensemble rewards help: Combining all reward functions gives the best overall results, 79.3 on Aircrafts-102 and 88.6 on Pets-37, while the format-plus-classification reward pair alone gives 76.3 and 86.8.
Methodology in Plain English
The authors start by testing whether adding a reasoning step changes answers. They evaluate several open-source MLLMs on four FGVC datasets under two prompting styles, answer-only and CoT, framing classification as an open-ended question-answering task.
They then train models with reinforcement fine-tuning, following the Visual-RFT setup, using a format reward (does the output use the required <think>...</think> and <answer>...</answer> tags), a binary classification reward (does the answer match the ground truth), and a thinking-length reward that gives 1 only if the reasoning length falls within a chosen interval. By varying that interval, they can directly control how much the model is allowed to reason and observe how accuracy responds. They also generate CoT training data with GPT-4o for supervised fine-tuning to check whether better reasoning text fixes the problem.
Based on the finding that shorter reasoning is better, they build ReFine-RFT. Its reward is an ensemble: the rule-based rewards above, plus an MLLM-based accuracy reward where a teacher model grades each prediction from 0 to 10 (normalized to [0, 1]) to handle semantically correct but lexically different answers, plus an embedding similarity reward using cosine similarity between the predicted and reference answer embeddings. Because these rewards differ in scale, convergence speed, and saturation point, summing them before normalization would let one reward dominate. MRN instead normalizes each reward within the sampled group of responses separately, then sums the normalized advantages before the policy update.
Training uses Qwen2-VL-2B-Instruct as the base model on four NVIDIA H100 GPUs with 81G of memory, Qwen2-VL-7B-Instruct as the MLLM judge, E5 as the embedding model, L_min=0 and L_max=10 for the length reward, LoRA applied to the attention and MLP projection modules, a learning rate of 2e-5, an accumulated batch size of 64, G=8 generations and β=0.04 for GRPO, a maximum of 200 training steps, and a completion length capped at 256 tokens.
Why This Matters
The paper reframes a common assumption: that more explicit reasoning is generally better. For perception-centric tasks, the amount of deliberation appears to be the dominant factor, which changes how practitioners should design prompts, training signals, and reward functions for multimodal models.
Real-world applications:
- Ecology and biodiversity monitoring: distinguishing plant varieties or species from images where categories differ only in fine patterns.
- Medical imaging: supporting subcategory-level diagnostic distinctions where subtle visual cues matter.
- Industrial inspection: identifying specific part or product variants that look nearly identical.
- Object-centric visual question answering: a model that cannot reliably separate similar categories such as pet breeds will also fail follow-up questions about the same subject.
Industry relevance: The results show a 2B backbone trained with 4-shot data and LoRA surpassing an 8B model trained on full FGVC datasets, which points toward cheaper, parameter-efficient adaptation for specialized fine-grained domains. MRN is described as plug-and-play, so it can be dropped into existing multi-reward RL pipelines, and the reasoning-length finding offers a simple lever for controlling inference cost and latency in deployed systems.
Future Directions
- Probing the underlying mechanism behind the "Cost of Thinking" — why longer textual reasoning harms fine-grained visual discrimination rather than refining it.
- Extending ReFine-RFT to broader multimodal tasks beyond FGVC.
- Determining how sensitive results are to the reasoning-length bounds (L_min, L_max) and whether an optimal range can be learned rather than set manually.
- Investigating the risk of reward hacking, since reasoning collapse also occurred without any explicit length constraint, and mitigating scoring bias in the MLLM-based accuracy reward, which the authors partially address with few-shot grading examples and a complementary embedding similarity reward.
Target Audience
Researchers and practitioners working on multimodal large language models, reinforcement learning from verifiable or multi-component rewards, and fine-grained visual recognition. It is most useful to readers already comfortable with GRPO, LoRA, and reward design, and to engineers who need parameter-efficient ways to adapt MLLMs to specialized visual domains with limited labeled data.
Authors’ abstract
Multi-modal large language models (MLLMs) exhibit strong general-purpose capabilities, yet still struggle on Fine-Grained Visual Classification (FGVC), a core perception task that requires subtle visual discrimination and is crucial for many real-world applications. A widely adopted strategy for boosting performance on challenging tasks such as math and coding is Chain-of-Thought (CoT) reasoning. However, several prior works have reported that CoT can actually harm performance on visual perception tasks. These studies, though, examine the issue from relatively narrow angles and leave open why CoT degrades perception-heavy performance. We systematically re-examine the role of CoT in FGVC through the lenses of zero-shot evaluation and multiple training paradigms. Across these settings, we uncover a central paradox: the degradation induced by CoT is largely driven by the reasoning length, in which longer textual reasoning consistently lowers classification accuracy. We term this phenomenon the ``Cost of Thinking''. Building on this finding, we make two key contributions: (1) MRN, a simple and general plug-and-play normalization method for multi-reward optimization that balances heterogeneous reward signals, and (2) ReFine-RFT, a framework that combines ensemble rewards with MRN to constrain reasoning length while providing dense accuracy-oriented feedback. Extensive experiments demonstrate the effectiveness of our findings and the proposed ReFine-RFT, achieving state-of-the-art performance across FGVC benchmarks. Project page: \href{https://refine-rft.github.io/}{ReFine-RFT}.