AI.info
Latest Research — Page 23
Browse Latest Research on AI.info.
- Video-KTR: Reinforcing Video Reasoning via Key Token Attribution
- Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs
- Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile Manipulation
- QuEPT: Quantized Elastic Precision Transformers with One-Shot Calibration for Multi-Bit Switching
- IB-GAN: Disentangled Representation Learning with Information Bottleneck Generative Adversarial Networks
- VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
- Context-Nav: Context-Driven Exploration and Viewpoint-Aware 3D Spatial Reasoning for Instance Navigation
- Dual-objective Language Models: Training Efficiency Without Overfitting
- Learning with less: label-efficient land cover classification at very high spatial resolution using self-supervised deep learning
- Continual Adaptation for Pacific Indigenous Speech Recognition
- Dexbotic: Open-Source Vision-Language-Action Toolbox
- Unified Spherical Frontend: Learning Rotation-Equivariant Representations of Spherical Images from Any Camera
- SLICEChat: Progressive In-Encoder Token Pruning for Whole-Slide Pathology Language Models
- CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia
- What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets
- BiKC+: Bimanual Hierarchical Imitation with Keypose-Conditioned Coordination-Aware Consistency Policies
- FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging
- When Does Predictor-Based RL Align with Human Perception? A Study of Subjective Rewards in Codec-Based Speech Language Models
- Fast Converging 3D Gaussian Splatting for 1-Minute Reconstruction
- Cornserve: A Distributed Serving System for Any-to-Any Multimodal Models
- Active Slice Discovery in Large Language Models
- BarcodeMamba+: Advancing State-Space Models for Fungal Biodiversity Research
- ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
- Explainable AI for the EU Right to Explanation: A Systematic Review of the Law-XAI Translation Gap
- InSight: A Benchmark for Agentic Claim Verification in Interactive Visualizations
- UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing
- VoxServe: Streaming-Centric Serving System for Speech Language Models
- "As Eastern Powers, I will veto." : An Investigation of Nation-level Bias of Large Language Models in International Relations
- BMAM: Brain-inspired Multi-Agent Memory Framework
- Multi-View Graph Learning with Graph-Tuple
- To Mix or To Merge: Toward Multi-Domain Reinforcement Learning for Large Language Models
- ContractScrub: A benchmark for final review of legal contracts
- Orient Anything V2: Unifying Orientation and Rotation Understanding
- CuMA: Aligning LLMs with Sparse Cultural Values via Demographic-Aware Mixture of Adapters
- Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception
- Revisiting Generative Infrared and Visible Image Fusion Based on Human Cognitive Laws
- From Internal Diagnosis to External Auditing: A VLM-Driven Paradigm for Data-Free Online Backdoor Defense
- Improving Flow Matching by Aligning Flow Divergence
- Kinship Data Benchmark for Multi-hop Reasoning
- SkyReels-Text: Fine-Grained Font-Controllable Text Editing for Poster Design