Skip to content
AI.info
News
Tools
Research
Learn
AI.info
Tianyi Lin
Explore Tianyi Lin on AI.info.
A Zeroth-Order Paradigm for LLM Preference Alignment
R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model?
Reward-free Alignment for Conflicting Objectives
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
Page 1