AI.info
Zhe Li
Explore Zhe Li on AI.info.
- CoCo-IR: Contextual Composed Image Retrieval
- Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token Interactions
- Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
- ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking
- MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection
- URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model
- BLM$_1$: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning