AI.info
Hao Li
Explore Hao Li on AI.info.
- Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs
- PhysInOne: Visual Physics Learning and Reasoning in One Suite
- DiverseDiT: Towards Diverse Representation Learning in Diffusion Transformers
- Chain of World: World Model Thinking in Latent Motion
- RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation
- Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction
- VAMOS-OCTA: Vessel-Aware Multi-Axis Orthogonal Supervision for Inpainting Motion-Corrupted OCT Angiography Volumes
- MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources
- Learning Sewing Patterns via Latent Flow Matching of Implicit Fields
- Agentic reinforcement learning empowers next-generation chemical language models for molecular design and synthesis
- Diff-MN: Diffusion Parameterized MoE-NCDE for Continuous Time Series Generation with Irregular Observations
- LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing
- TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers
- Towards Generative Location Awareness for Disaster Response: A Probabilistic Cross-view Geolocalization Approach
- RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
- FDP: A Frequency-Decomposition Preprocessing Pipeline for Unsupervised Anomaly Detection in Brain MRI
- Agentmandering: A Game-Theoretic Framework for Fair Redistricting via Large Language Model Agents
- Monocular absolute depth estimation from endoscopy via domain-invariant feature learning and latent consistency
- IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction
- From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
- InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
- ICL-Router: In-Context Learned Model Representations for LLM Routing
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints