AI.info
Latest Research — Page 118
Browse Latest Research on AI.info.
- Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation
- Task adaptation of Vision-Language-Action model: 1st Place Solution for the 2025 BEHAVIOR Challenge
- Lifelong Domain Adaptive 3D Human Pose Estimation
- Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalization
- Enhancing Multimodal Protein Function Prediction Through Dual-Branch Dynamic Selection with Reconstructive Pre-Training
- Thinking in Structures: Evaluating Spatial Intelligence in Constraint-Governed Spaces
- PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
- FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation
- COOPERTRIM: Adaptive Data Selection for Uncertainty-Aware Cooperative Perception
- Ctrl-F-Resist. Practices, Challenges, and Technical Needs of Civil Society Organizations Monitoring the Far-Right Online
- Towards participatory speech dataset curation: A queer case study and conceptual framework
- When Agents "Misremember" Collectively: Exploring the Mandela Effect in LLM-based Multi-Agent Systems
- LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference
- MicroProbe: Efficient Reliability Assessment for Foundation Models with Minimal Data
- Scaling Vision Transformers for Functional MRI with Flat Maps
- Probability-Entropy Calibration: An Elastic Indicator for Adaptive Fine-tuning
- Enhancing Procedural Writing Through Personalized Example Retrieval: A Case Study on Cooking Recipes
- Manipulation of Deformable Linear Objects Using Model Predictive Path Integral Control with Bidirectional Long Short-Term Memory Learning
- SEMA: Simple yet Effective Learning for Multi-Turn Jailbreak Attacks
- MMEarth-Bench: Global Model Adaptation via Multimodal Test-Time Training
- SMG: Semantic Motion Graph for Monocular Dynamic Gaussian Splatting
- VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models
- Do What I Say: A Spoken Prompt Dataset for Instruction-Following
- AcademicEval: Live Long-Context LLM Benchmark
- Deployment-Time Memorization in Foundation-Model Agents
- Focusing on Language: Revealing and Exploiting Language Attention Heads in Multilingual Large Language Models
- MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
- SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
- On the Hardness of Conditional Independence Testing In Practice
- Distilling the Thought, Watermarking the Answer: A Principle Semantic Guided Watermark for Large Reasoning Models
- Personalized Image Editing in Text-to-Image Diffusion Models via Collaborative Direct Preference Optimization
- A Space-Time Transformer for Precipitation Nowcasting
- Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders
- VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
- ORGAN: Object-Centric Representation Learning using Cycle Consistent Generative Adversarial Networks
- Towards Inference-time Scaling for Continuous Space Reasoning
- T3former: Temporal Graph Classification with Topological Machine Learning
- SeeU: Seeing the Unseen World via 4D Dynamics-aware Generation
- IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
- Multimodal Assessment of Pancreatic Cancer Resectability Using Deep Learning