Research
PreferThinker: Reasoning-based Personalized Image Preference Assessment
PreferThinker: Reasoning-based Personalized Image Preference Assessment Overview Research area: Multimodal large language models (MLLMs), reinforcement-learned reasoning, and human-preference / image-
- arXiv
- 2511.00609
- Published
- 2025-11-01
- Authors
- Shengqi Xu, Xinpeng Zhou, Yabo Zhang, Ming Liu, Tao Liang, Tianyu Zhang, Yalong Bai, Zuxuan Wu, Wangmeng Zuo
AI summary
PreferThinker: Reasoning-based Personalized Image Preference AssessmentOverview
Research area: Multimodal large language models (MLLMs), reinforcement-learned reasoning, and human-preference / image-quality assessment — specifically the sub-problem of personalized image preference assessment.
Technical level: Advanced. The paper assumes familiarity with vision-language model fine-tuning, Chain-of-Thought (CoT) supervision, and GRPO-style reinforcement learning from verifiable rewards.
One-sentence scope: The paper proposes PreferThinker, a reasoning-based system that predicts an individual user's visual "preference profile" from a few reference images and then uses that profile to produce interpretable, multi-dimensional assessments of candidate images.
Paper metadata: arXiv:2511.00609 (v4, dated 26 Jun 2026 in the paper footer), licensed CC BY 4.0. Authors are affiliated with Fudan University, Harbin Institute of Technology, and iN2X.
What This Paper Is About
Most image-preference models score general qualities such as text-image alignment and aesthetics, trained on large-scale data where most users share similar criteria. Personalized assessment is much harder: each user's data is limited, not easily scalable, and individual tastes are complex and diverse (art style, color, medium, and so on). PreferThinker's goal is to assess a specific user's preferences using only a small set of that user's reference images (both preferred and non-preferred) plus two candidate images, and to do so in a way that is both accurate and interpretable.
Key Contributions
- PreferThinker system. A reasoning-based personalized image assessment system with preference profile prediction that follows a "predict-then-assess" CoT structure: it predicts the user's preference and non-preference profiles from reference images, then scores candidates across multiple dimensions with interpretable reasoning grounded in those profiles.
- PreferImg-CoT dataset. Described as the first CoT-style personalized assessment dataset, annotated with preference profiles and high-quality CoT-style reasoning, comprising 60,000 user samples and built on a larger generated dataset, PreferImg (80K simulated users, of whom 20K are associated with multiple preference profiles, plus 1.36 million images).
- Two-stage training plus a new reward. Cold-start supervised fine-tuning (SFT) to elicit structured reasoning, followed by GRPO-based reinforcement fine-tuning (RFT) for exploration and generalization, together with a novel similarity-aware prediction reward that rewards accurate preference-profile prediction.
- Empirical validation. Extensive comparisons against CLIP-based, MLLM-based, open-source, and closed-source baselines show PreferThinker outperforming prior state of the art, and the predicted profile is shown to also benefit personalized image generation.
Main Findings
- Top accuracy on the proposed benchmark. On the 1,500-user PreferImg benchmark, PreferThinker (7B parameters) reaches 96.6 (seen, single-preference), 92.0 (seen, multi-preference), 96.4 (unseen, single-preference) and 92.8 (unseen, multi-preference), with an average of 88.7. That is +2.4, +8.8, +3.0, +6.8 over the previous state of the art and +21.2, +30.0, +24.4, +28.0 over its base model (Qwen2.5-VL-7B).
- Closing the gap with closed-source models. Claude 3.7 (93.8 / 83.2 / 90.2 / 86.0; average 83.6) and OpenAI GPT-4o (94.2 / 80.4 / 92.2 / 85.2; average 83.5) are strong but still below PreferThinker overall, as are Doubao-1.5-vision-pro-32k (average 80.2) and Gemini-2.5-flash (average 71.7).
- General open MLLMs and CLIP-based scorers struggle. CLIP-based models (PickScore 54.7, ImageReward 54.1, HPSv2 54.6, CLIPScore 53.4, Aesthetics 49.3, CycleReward 53.5 average) and MLLM-based reward models (UnifiedReward 52.9, UnifiedReward-Think 51.4, LLaVA-Reward 55.6) perform near or below chance on personalized assessment. ViPer reaches 81.2 average but drops notably on multi-preference data (78.0 seen, 80.0 unseen) and outputs only a numerical score.
- Multi-preference data is the hardest setting. Most baselines degrade sharply from single- to multi-preference users; ViPer falls from 92.4 to 78.0 on seen data, while PreferThinker falls only from 96.6 to 92.0.
- Real-user benchmark is harder and general-preference flavored. On PickaPic, reformatted by user ID into 894 user samples, PreferThinker scores 65.7, tied with GPT-4o and second to PickScore's 67.9 — PickScore benefits because PickaPic's labels still reflect general preferences and serve as seen data for it. The paper reports a -2.2 gap versus the previous state of the art on this column.
- Both training stages matter. Ablations show base-only performance of 75.4 / 72.0 / 62.0 / 64.8 (seen-SP, unseen-SP, seen-MP, unseen-MP assessment accuracy). SFT alone raises this to 92.0 / 91.8 / 81.2 / 81.6; SFT + RL reaches 94.9 / 95.2 / 90.4 / 89.2; adding the prediction reward reaches 96.6 / 96.4 / 92.0 / 92.8. RL after SFT especially improves unseen multi-preference users.
- The prediction reward improves profile prediction. With the prediction reward, profile-prediction accuracy rises to 85.4 / 86.2 / 78.6 / 80.7 across the four settings, versus 84.1 / 85.2 / 72.5 / 74.1 without it. The paper links better profile prediction to more reasonable assessments and correct final answers.
- Reasoning can also be used implicitly and quickly. Without explicit CoT, PreferThinker still scores 93.2 on PreferImg and 65.2 on PickaPic while being the fastest at 0.80 s, compared with ViPer (86.0 / 62.2, 2.19 s), Doubao-1.5-vision (84.3 / 63.8, 5.39 s), GPT-4o (88.0 / 65.7, 12.92 s) and Claude 3.7 (88.3 / 64.9, 18.22 s).
- Robust to few reference images. Across different numbers of personalized reference images, PreferThinker outperforms other baselines on both seen and unseen data.
- Profiles transfer to generation. Optimizing initial text prompts with the predicted preference profile produces generated images that align well with the reference images.
Methodology in Plain English
The core insight is that while every user's taste is unique, the elements that make up that taste are shared. The authors identify 15 visual elements from Lexica text prompts and ran a user study with 100 participants, asking each to pick the five most important ones. The top choices were art style, color, detail, art medium and saturation. They collected a vocabulary of 288 related terms across these elements to keep profiles diverse.
Using that vocabulary, they simulated 80,000 users, each assigned a distinct preference profile (with multiple profiles assigned to some users), and generated preferred and non-preferred reference images plus two candidate images for each, drawing on 190K prompts from Lexica, DiffusionDB and COCO — producing 1.36 million images. The advanced Claude 3.7 model was prompted to write CoT-style reasoning over these samples, following a predict-then-assess template (profile prediction, multi-dimensional scoring, final answer), and illogical or inconsistent responses were filtered out, yielding PreferImg-CoT with 60,000 user samples.
Training uses Qwen2.5-VL-7B as the base. Stage one is cold-start SFT with a standard token-level cross-entropy loss, teaching the model the structured reasoning format. Stage two applies GRPO reinforcement learning, in which the model generates a group of candidate responses whose scalar rewards are normalized into group-relative advantages and optimized without a critic model. Rewards are a weighted sum of three terms: a similarity-aware prediction reward (text similarity between predicted and ground-truth profiles measured with SBERT, plus image similarity between images generated from predicted versus ground-truth profiles, measured with DreamSim), a format correctness reward (1 if profile tags, think tags and answer tags are correctly used), and an assessment accuracy reward (1 if the predicted answer matches ground truth). Weights are 0.7, 0.3 and 1.0 respectively. Training used 8 NVIDIA A100 GPUs, LLaMA-Factory for SFT (one epoch), the VLM-R1 framework for RL, a learning rate of 1e-5 and 6 rollout samples per input.
Why This Matters
Research impact. The work reframes personalized assessment as a profile-prediction problem, arguing that shared visual elements let scarce per-user data be supplemented by large-scale simulated data. It is presented as the first reasoning-based personalized assessment system, and it introduces a hand-designed reward that couples text and image similarity for supervision where only discrete answers exist. It also provides a public-facing dataset contribution in an area where the closest prior personalized dataset was simulated and unreleased.
Real-world applications:
- Personalized content recommendation, where ranking must reflect an individual's taste rather than a population average.
- Evaluating and aligning generative models with individual human preferences, beyond generic aesthetics or text-image alignment scores.
- Personalized text-to-image generation, using the predicted preference profile to steer prompts (demonstrated in the paper).
- Interpretable, multi-dimensional feedback on candidate images, giving users or designers a breakdown of why an image fits or does not fit their stated preferences, rather than a single opaque number.
Industry relevance. The system is built on a 7B open base model and runs at 0.80 s without explicit CoT, which is faster than the commercial and prior open baselines measured. That combination of high accuracy, interpretability and latency is relevant to recommendation feeds, creative tools, and preference-data pipelines where per-user interaction data is too thin to train dedicated models.
Future Directions
- Real-user data. The paper notes no real-user dataset exists for personalized preference assessment, so evaluation relies on simulated users plus a repurposed general-preference dataset (PickaPic, 894 user samples, labels still reflecting general preferences). Collecting genuine personalized labels is an open need.
- Closing the real-world gap. PreferThinker trails PickScore by 2.2 points on PickaPic; reconciling profile-based personalized reasoning with datasets labeled by general preference remains unresolved.
- Scaling and profile coverage. The current profile uses 15 visual elements and 288 terms; whether the framework extends to finer or other modalities of preference is not reported.
- Limitations and future work. The paper states that limitations and future work are discussed in Section E.4 of the appendix, along with transferability to single-image assessment and further generation benefits (Sections E.2 and
Authors’ abstract
Personalized image preference assessment aims to evaluate an individual user's image preferences by relying only on a small set of reference images as prior information. Existing methods mainly focus on general preference assessment, training models with large-scale data to tackle well-defined tasks such as text-image alignment. However, these approaches struggle to handle personalized preference because user-specific data are scarce and not easily scalable, and individual tastes are often diverse and complex. To overcome these challenges, we introduce a common preference profile that serves as a bridge across users, allowing large-scale user data to be leveraged for training profile prediction and capturing complex personalized preferences. Building on this idea, we propose a reasoning-based personalized image preference assessment framework that follows a \textit{predict-then-assess} paradigm: it first predicts a user's preference profile from reference images, and then provides interpretable, multi-dimensional scores and assessments of candidate images based on the predicted profile. To support this, we first construct a large-scale Chain-of-Thought (CoT)-style personalized assessment dataset annotated with diverse user preference profiles and high-quality CoT-style reasoning, enabling explicit supervision of structured reasoning. Next, we adopt a two-stage training strategy: a cold-start supervised fine-tuning phase to empower the model with structured reasoning capabilities, followed by reinforcement learning to incentivize the model to explore more reasonable assessment paths and enhance generalization. Furthermore, we propose a similarity-aware prediction reward to encourage better prediction of the user's preference profile, which facilitates more reasonable assessments exploration. Extensive experiments demonstrate the superiority of the proposed method.