Skip to content
AI.info

Research

Measurement Plasticity: Sensor-Level Adaptation for Vision-Language Models

Overview Research area: Computer vision and machine learning — specifically test-time adaptation (TTA) and robustness of vision-language models (VLMs) under sensor-level distribution shift. Technical

Measurement Plasticity: Sensor-Level Adaptation for Vision-Language Models
arXiv
2512.12571
Published
2025-12-14
Authors
Boyeong Im, Wooseok Lee, Yoojin Kwon, Hyung-Sin Kim

AI summary

Overview

  • Research area: Computer vision and machine learning — specifically test-time adaptation (TTA) and robustness of vision-language models (VLMs) under sensor-level distribution shift.
  • Technical level: Advanced. The paper assumes familiarity with CLIP-style VLMs, prompt tuning, test-time adaptation, feature statistics (mean/variance of visual tokens), and camera exposure parameters.
  • Scope: A single framework paper introducing "measurement plasticity" — moving test-time adaptation upstream from model tokens to the physical measurement process — evaluated on two controlled sensor

Authors’ abstract

We propose Multi-View Physical-prompt (MVP) for Test-Time Adaptation (TTA), a forward-only framework that moves TTA from tokens to photons by treating the camera exposure triangle (i.e., ISO, shutter speed, and aperture) as physical prompts. At inference, MVP acquires selected multiple physical views using a source-affinity score, evaluates digitally augmented variants of each retained view and filters the lowest-entropy predictions, and aggregates predictions with hard voting. This selection-then-vote design is simple, calibration-friendly, and requires no gradients or model modifications. On ImageNet-ES and ImageNet-ES-Diverse, MVP outperforms digital-only TTA on both Auto-Exposure and a combination with conventional sensor control. MVP remains effective under reduced parameter candidates that lower capture latency, demonstrating its practicality.

Read the original paper