AI.info
Haitao Mi
Explore Haitao Mi on AI.info.
- Verified Critical Step Optimization for LLM Agents
- Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification
- Stable and Efficient Single-Rollout RL for Multimodal Reasoning
- DeepCompress: A Dual Reward Strategy for Dynamically Exploring and Compressing Reasoning Chains