AI.info
Latest Research — Page 37
Browse Latest Research on AI.info.
- EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs
- From Linear to Nonlinear: Provable Weak-to-Strong Generalization through Feature Learning
- UNet-Based Keypoint Regression for 3D Cone Localization in Autonomous Racing
- Eye-Tracking, Mouse Tracking, Stimulus Tracking,and Decision-Making Datasets in Digital Pathology
- YOLO-IOD: Towards Real Time Incremental Object Detection
- Randomization Boosts KV Caching, Learning Balances Query Load: A Joint Perspective
- Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans
- CloDS: Visual-Only Unsupervised Cloth Dynamics Learning in Unknown Conditions
- TransBiolab: A Real-World Multi-View Dataset of Cluttered Transparent Biomedical Objects
- Liars' Bench: Evaluating Lie Detectors for Language Models
- Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search
- When Token Pruning is Worse than Random: Understanding Visual Token Information in VLLMs
- NADIR: Differential Attention Flow for Non-Autoregressive Transliteration in Indic Languages
- TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
- EMFusion: Uncertainty-Aware Conditional Diffusion Model for Multivariate Narrow-band Exposure Forecasting
- EPSegFZ: Efficient Point Cloud Semantic Segmentation for Few- and Zero-Shot Scenarios with Language Guidance
- A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans
- TokEval: A Tokenizer Evaluation Suite
- ChatR1: Reinforcement Learning for Conversational Reasoning and Retrieval Augmented Question Answering
- TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
- DADO: A Depth-Attention framework for Object Discovery
- RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba
- EmoMed: An Emotionally-Aware Agent for Multimodal Medical Support with Real-Time Information Retrieval
- Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation
- Resisting Manipulative Bots in Meme Coin Copy Trading: A Multi-Agent Approach with Chain-of-Thought Reasoning
- Towards Detecting AI-Assisted Responses in Online Surveys
- VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models
- SpeakerSleuth: Can Large Audio-Language Models Judge Speaker Consistency across Multi-turn Dialogues?
- POLARIS: Typed Planning and Governed Execution for Agentic AI in Back-Office Automation
- PRInTS: Reward Modeling for Long-Horizon Information Seeking
- NeuroRefiner: Morphology-Aware Multi-Agent Refinement for 3D Fluorescence Microscopy Neuron Segmentation
- Towards Robust Pseudo-Label Learning in Semantic Segmentation: An Encoding Perspective
- Envisioning the Future, One Step at a Time
- TreeQ: Pushing the Quantization Boundary of Diffusion Transformer via Tree-Structured Mixed-Precision Search
- d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models
- SimDiff: Simpler Yet Better Diffusion Model for Time Series Point Forecasting
- Unified all-atom molecule generation with neural fields
- Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning
- Proxy Compression for Language Modeling
- ROBoto2: An Interactive System and Dataset for LLM-assisted Clinical Trial Risk of Bias Assessment