AI.info
Latest Research — Page 3
Browse Latest Research on AI.info.
- Auto-Regressive Masked Diffusion Models
- MixRI: Mixing Features of Reference Images for Novel Object Pose Estimation
- Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
- Mind-to-Face: Neural-Driven Photorealistic Avatar Synthesis via EEG Decoding
- QuCo-RAG: Quantifying Uncertainty from the Pre-training Corpus for Dynamic Retrieval-Augmented Generation
- Adaptive Confidence Regularization for Multimodal Failure Detection
- Understanding the Effects of Distractors on Reasoning Vision-Language Models
- Know your Trajectory -- Trustworthy Reinforcement Learning deployment through Importance-Based Trajectory Analysis
- The False Promise of Zero-Shot Super-Resolution in Machine-Learned Operators
- Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views
- Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models
- GraspLDP: Towards Generalizable Grasping Policy via Latent Diffusion
- Enhancing Generalization of Depth Estimation Foundation Model via Weakly-Supervised Adaptation with Regularization
- Multimodal Bandits: Regret Lower Bounds and Optimal Algorithms
- CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations
- CoFEH: LLM-driven Feature Engineering Empowered by Collaborative Bayesian Hyperparameter Optimization
- DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
- Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following
- Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
- Anatomy Contextualized Adaption of CT Foundation Models
- Coarse-Grained Boltzmann Generators
- WED-Net: A Weather-Effect Disentanglement Network with Causal Augmentation for Urban Flow Prediction
- Integrating Multi-view Multi-light Surface Reconstruction into Cultural Heritage Workflows
- Zero-RAG: Towards Retrieval-Augmented Generation with Zero Redundant Knowledge
- Can Local Learning Match Self-Supervised Backpropagation?
- SPDMark: Selective Parameter Displacement for Robust Video Watermarking
- Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot Learning
- TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards
- AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance
- Saudi Sign Language Translation Using T5
- LLMs for Game Theory: Entropy-Guided In-Context Learning and Adaptive CoT Reasoning
- InterMoE: Individual-Specific 3D Human Interaction Generation via Dynamic Temporal-Selective MoE
- When One Modality Rules Them All: Backdoor Modality Collapse in Multimodal Diffusion Models
- Rethinking Language's Role in Efficient VLA for Autonomous Vehicles: Toward Smarter, Trustworthy Driving
- LogicSkills: A Structured Benchmark for Formal Reasoning in Large Language Models
- Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
- Overview of CHIP 2025 Shared Task 2: Discharge Medication Recommendation for Metabolic Diseases Based on Chinese Electronic Health Records
- Spark: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic Learning
- Nonlinear Model Predictive Control of a Robotic Soft Esophagus
- MetaCaDI: A Meta-Learning Framework for Causal Discovery from Multiple Environments with Unknown Interventions