AI.info
Latest Research — Page 49
Browse Latest Research on AI.info.
- Pass the Baton: Trajectory-Relayed On-Policy Distillation
- BulletTime: Decoupled Control of Time and Camera Pose for Video Generation
- LumiCtrl : Learning Illuminant Prompts for Lighting Control in Personalized Text-to-Image Models
- AOMGen: Photoreal, Physics-Consistent Demonstration Generation for Articulated Object Manipulation
- DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
- LoRAverse: A Submodular Framework to Retrieve Diverse Adapters for Diffusion Models
- BitMar: Low-Bit Multimodal Fusion with Episodic Memory for Edge Devices
- Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders
- DictPFL: Efficient and Private Federated Learning on Encrypted Gradients
- Logit-Based Losses Limit the Effectiveness of Feature Knowledge Distillation
- Learning to Make Friends: Coaching LLM Agents toward Emergent Social Ties
- Stop Wasting Your Tokens: Towards Efficient Runtime Multi-Agent Systems
- Understanding and Improving Hyperbolic Deep Reinforcement Learning
- MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression Segmentation
- A Unified Detection Pipeline for Robust Object Detection in Fisheye-Based Traffic Surveillance
- Egocentric Co-Pilot: Web-Native Smart-Glasses Agents for Assistive Egocentric AI
- Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
- CodeDistiller: Automatically Generating Code Libraries for Scientific Coding Agents
- BuildOcc: A Large Language Model Occupant Agent Platform for Building Energy Research
- Vision Transformer Finetuning Benefits from Non-Smooth Components
- Head Pursuit: Probing Attention Specialization in Multimodal Transformers
- Learning visual representations for compositional analysis of artworks and photographs
- MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning
- Wait, Wait, Wait... Why Do Reasoning Models Loop?
- Distilling Feedback into Memory-as-a-Tool
- DuoLens: A Framework for Robust Detection of Machine-Generated Multilingual Text and Code
- VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset Distillation
- Resp-Agent: An Agent-Based System for Multimodal Respiratory Sound Generation and Disease Diagnosis
- ELMO: Efficiency via Low-precision and Peak Memory Optimization in Large Output Spaces
- Towards Long-Horizon Interpretability: Efficient and Faithful Multi-Token Attribution for Reasoning LLMs
- XAI-MeD: Explainable Knowledge Guided Neuro-Symbolic Framework for Domain Generalization and Rare Class Detection in Medical Imaging
- Discretizing Continuous Time Series for Imputation with Masked Diffusion Training
- Learning Sparse Approximate Inverse Preconditioners for Conjugate Gradient Solvers on GPUs
- QiMeng-NeuComBack: Self-Evolving Translation from IR to Assembly Code
- Precise Top-Layer Fabric Segmentation for Fabric Destacking with Edge- and Shape-Aware Deep Networks
- A Comprehensive Dataset for Human vs. AI Generated Text Detection
- Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents
- Pixel Super-Resolved Fluorescence Lifetime Imaging Using Deep Learning
- Interpretable Vision Transformers in Monocular Depth Estimation via SVDA