Research
Understanding Task Transfer in Vision-Language Models
Overview Research area: Multimodal machine learning, specifically vision-language models (VLMs), visual perception, and transfer learning. Technical level: Intermediate. The core ideas are accessible,

- arXiv
- 2511.18787
- Published
- 2025-11-24
- Authors
- Bhuvan Sachdeva, Karan Uppal, Abhinav Java, Vineeth N. Balasubramanian
AI summary
Overview
- Research area: Multimodal machine learning, specifically vision-language models (VLMs), visual perception, and transfer learning.
- Technical level: Intermediate. The core ideas are accessible, but familiarity with LoRA finetuning, zero-shot evaluation, and benchmark terminology helps.
- Scope: A systematic empirical study of how finetuning a VLM on one visual perception task changes its zero-shot performance on 13 other perception tasks, using a newly proposed normalized metric called the Perfection Gap Factor (PGF).
What This Paper Is About
VLMs score well on multimodal benchmarks but still trail humans and specialized models on basic visual perception tasks such as depth estimation or object counting. Practitioners often finetune VLMs on task-specific data to close this gap, but nobody had systematically measured how such finetuning affects the model's performance on other perception tasks. This paper asks that question directly and quantifies the answer for three open-weight Qwen-2.5-VL models across 13 BLINK perception tasks.
Key Contributions
- Systematic study of perception task transfer: The authors state they are the first to analyze how finetuning on one visual perception task affects zero-shot performance on a broad suite of other perception tasks.
- Perfection Gap Factor (PGF): A new metric that measures how much of the remaining gap to a task's ceiling is closed or opened by finetuning on a source task, normalizing for heterogeneous task difficulty and model baselines.
- Task properties and structure: Identification of scale-dependent sharpening of transfer, transfer behavior tied to perceptual level and granularity, task cliques of mutually beneficial or detrimental tasks, and "task personas" (Donors, Pirates, Sponges, Sieves).
- Beyond images and downstream use of PGF: Cross-task transfer is also evaluated on video (spatio-temporal) tasks, and PGF is used to select training data subsets that improve finetuning efficiency while mitigating negative transfer.
Main Findings
- Low-level tasks transfer best: Relative Depth, Relative Reflectance, and Visual Correspondence are the most positively transferable and the most malleable tasks; the authors conclude finetuning on low-level tasks is beneficial compared with mid- and high-level tasks.
- Granularity matters: Image-level tasks (Art Style, Counting, Forensic Detection, and others) show higher positive transferability than pixel- and crop-level tasks. Both image-level and pixel-level tasks show higher positive malleability than crop-level tasks.
- Scale sharpens transfer: Average positive transferability increases with model size across Qwen-2.5-VL 3B, 7B, and 32B. There is no consistent trend for average negative transferability. Positive malleability also increases with model size.
- Negative transfer is not uniform: Negative transferability and malleability show a sharp spike in Qwen2.5-VL-7B. Averaged across models, low-level and image-level tasks show the highest magnitude of negative transferability, while high-level and crop-level tasks show the highest magnitude of negative malleability.
- Cliques of cooperation and interference: In the 3B and 7B variants the largest cliques comprise 3 to 4 tasks, while Qwen-2.5-VL 32B contains a maximal positive clique of size 9: {Art Style, Counting, Functional Correspondence, Jigsaw, Relative Depth, Relative Reflectance, Spatial Reasoning, Visual Correspondence, Visual Similarity}. A negative clique of size 4 is reported for the 32B model: {Forensics Detection, Multi-view Reasoning, Object Localization, Semantic Correspondence}.
- Task personas: Semantic Correspondence is identified as a statistically significant Donor task (p < 0.01 across models) and Functional Correspondence as a significant Pirate task in the 3B and 7B variants (p < 0.05). Visual Similarity, Relative Depth, and Relative Reflectance are Sponge tasks, with Visual Similarity and Relative Depth statistically significant across models (p < 0.001). Forensic Detection is the sole Sieve task, significant in the 3B and 32B variants (p < 0.005).
- Transfer generalizes to video: Evaluated on VSI Bench with Qwen-2.5-VL 3B and 7B, Relative Reflectance emerges as a donor and Forensic Detection as a pirate, matching the image results. Object Counting acts as a sponge, while Object Appearance Order and Object Relative Distance act as sieves.
- PGF-guided data selection works: Using Qwen-2.5-VL 7B, PGF-informed dataset mixtures consistently outperform randomly sampled mixtures, and even surpass direct finetuning on the target task in the cases of Jigsaw and Object Localisation. The authors describe this as a preliminary finding.
- PGF has asymmetric bounds: Positive PGF is capped at 1, achieved when finetuning fully closes the gap to perfection. Negative PGF has a finite lower bound of -(m-1) for a task with m evaluation questions; with m = 200 questions, PGF_min = -199. This asymmetry motivates separate reporting of positive and negative transferability.
Methodology in Plain English
The researchers took three open-weight Qwen-2.5-VL models (3B, 7B, 32B) and finetuned each one separately on each of 13 perception tasks from the BLINK benchmark using LoRA. Because BLINK only provides validation and test splits, they rebuilt training data by retrieving the original source datasets (for example TallyQA for Counting, Depth in the Wild for Relative Depth, Spair-71k for Semantic Correspondence, WikiArt for Art Style) while keeping BLINK's task definitions and response formats. Each finetuned checkpoint was then evaluated zero-shot on the validation splits of all 13 tasks, with every experiment repeated over four different random seeds.
To compare changes across tasks with very different difficulty levels, they introduced the Perfection Gap Factor: the accuracy change after finetuning divided by the remaining gap to the ceiling (with the ceiling set to 100 and a small constant of 10⁻⁶ added for numerical stability). From these PGF values they computed transferability (how broadly and strongly a source task influences others) and malleability (how sensitive a target task is to finetuning on other tasks), each split into positive and negative components with an exponential weighting that penalizes effects concentrated on only a few tasks. They also built task transfer graphs by keeping the strongest 20 percent of positive and negative PGF edges, used Wilcoxon tests to check clique stability across seeds and unpaired t-tests to validate personas, and ran a data-selection experiment on the 7B model comparing PGF-informed mixtures against random mixtures. Finetuning used 8xA100 40GB GPUs with DeepSpeed ZeRO-2 (3B and 7B) or ZeRO-3 (32B), batch size 16, weight decay 0, warmup ratio 0.03, cosine decay, LoRA rank 8, and alpha 16; GPT-4.1 was used to extract responses and evaluation used the official BLINK benchmark code.
Why This Matters
For research, the paper reframes perception tasks in VLMs as an interacting system rather than a set of independent capabilities, and it supplies a normalized metric that makes transfer measurements comparable across tasks with different baselines and ceilings. It also connects modern foundation-model behavior to the older transfer-learning literature (such as Taskonomy) that predates the foundation-model era.
Real-world applications implied by the tasks and findings:
- Data-efficient finetuning pipelines: When labeled data for a target task is unavailable, PGF can indicate which related datasets to mix in, potentially removing the need for direct supervision in some cases.
- Perception-critical deployments: Tasks such as relative depth, object counting, and spatial reasoning underpin robotics, navigation, and augmented reality, where knowing which finetuning jobs help or hurt matters.
- Forensics and document analysis: The identification of Forensic Detection as a Sieve task warns that finetuning on other perception tasks can degrade this capability.
- Multimodal video and embodied agents: The video results show that image-level perception finetuning transfers to spatio-temporal tasks like route planning and object appearance order.
For industry, the practical message is that task-specific finetuning of VLMs carries measurable risk of negative interference, and that PGF offers a cheap way to plan finetuning order and data mixtures, potentially improving results while avoiding performance regressions.
Future Directions
- Open-ended generation: The current analysis relies mainly on multiple-choice benchmarks, which restrict the output space and may suppress failure modes and transfer patterns that appear in open-ended generation.
- Newer and more diverse models: Extending the study to newer architectures would clarify how generalizable these findings are and how transfer behavior evolves with model development.
- Fuller study of PGF-guided data mixtures: The data-selection experiment is explicitly described as preliminary and limited to Qwen-2.5-VL 7B; the authors state that a comprehensive study is out of scope for this work and will be pursued in future research.
- Exploring negative cliques and hierarchical design: The framework enables discovery of negative cliques, and the authors suggest the low-level task results support the hypothesis that VLMs could benefit from hierarchical visual processing pathways.
Target Audience
Researchers and engineers working on vision-language models, multimodal finetuning, and transfer learning, particularly those who design LoRA-based finetuning pipelines or decide which perception datasets to train on. It is also useful for benchmark designers and practitioners evaluating multimodal models, since the metric definitions and task taxonomy are directly reusable. Readers without background in VLM finetuning will still follow the high-level findings but will need to look up terms such as LoRA, zero-shot evaluation, and PGF to engage with the details.
Authors’ abstract
Vision-Language Models (VLMs) perform well on multimodal benchmarks but lag behind humans and specialized models on visual perception tasks like depth estimation or object counting. Finetuning on one task can unpredictably affect performance on others, making task-specific finetuning challenging. In this paper, we address this challenge through a systematic study of task transferability. We examine how finetuning a VLM on one perception task affects its zero-shot performance on others. We introduce Perfection Gap Factor (PGF), a normalized metric that measures change in performance as a result of task transfer. We utilize PGF to compute Task Transferability, which captures both the breadth and the magnitude of transfer induced by a source task. Using three open-weight VLMs evaluated across 13 perception tasks, we construct a task transfer graph that reveals previously unobserved relationships among perception tasks. Our analysis uncovers patterns of positive and negative transfer, identifies groups of tasks that mutually influence each other, organizes tasks into personas based on their transfer behavior and demonstrates how PGF can guide data selection for more efficient training. These findings highlight both opportunities for positive transfer and risks of negative interference, offering actionable guidance for advancing VLMs.