AI.info
Latest Research — Page 22
Browse Latest Research on AI.info.
- PragWorld: A Benchmark Evaluating LLMs' Local World Model under Minimal Linguistic Alterations and Conversational Dynamics
- Motion4D: Learning 3D-Consistent Motion and Semantics for 4D Scene Understanding
- X-MULTI: VLM-based Imaging Factor Disentanglement for Factor-Aware Image Synthesis
- InstructMesh: Selective Refinement of Generative 3D Models for Fabrication
- From Knowledge to Treatment: Large Language Model Assisted Biomedical Concept Representation for Drug Repurposing
- Towards Efficient Evaluation of Evolutionary Transfer Optimization: Case Studies on Task-Parameterized Applications
- Empowering Medical Equipment Sustainability in Low-Resource Settings: An AI-Powered Diagnostic and Support Platform for Biomedical Technicians
- Reinforcement Learning with Backtracking Feedback
- DiA-gnostic VLVAE: Disentangled Alignment-Constrained Vision Language Variational AutoEncoder for Robust Radiology Reporting with Missing Modalities
- Clinician-in-the-Loop Smart Home System to Detect Urinary Tract Infection Flare-Ups via Uncertainty-Aware Decision Support
- CellPath-Bench: A Multidimensional Benchmark for Whole-Slide Cellular Representations in Pathology Foundation Models
- DriveLaW:Unifying Planning and Video Generation in a Latent Driving World
- Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
- Sparse Orthogonal Regression Technique: A Spectral Framework for Equation Discovery, Approximation, and Integration
- Differential 6-DOF Pose Estimation with Provable First-Order Immunity to Camera Calibration Errors
- Self-Evaluating LLMs for Multi-Step Tasks: Stepwise Confidence Estimation for Failure Detection
- Instant Personalized Large Language Model Adaptation via Hypernetwork
- DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation
- Discovering Differences in Strategic Behavior Between Humans and LLMs
- Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
- GeoDiff: Geometry-Guided Diffusion for Metric Depth Estimation
- GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks
- Stress-Testing Emotional Support Models: Moving from Homogeneous to Diverse Help Seekers
- Text2Loc++: Generalizing 3D Point Cloud Localization from Natural Language
- Unsupervised Ensemble Learning Through Deep Energy-based Models
- Finding the Sweet Spot: Trading Quality, Cost, and Speed During Inference-Time LLM Reflection
- Learning Self-Correction in Vision-Language Models via Rollout Augmentation
- MARS: Multi-Specialist LLM Relay System for Competitive Programming
- Ambient Dataloops: Generative Models for Dataset Refinement
- AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models
- Learning Ordinal Probabilistic Reward from Preferences
- Monitor-Generate-Verify (MGV): Formalising Metacognitive Theory for Language Model Reasoning
- Travel Time Prediction from Sparse Open Data
- Multimodal Classification via Total Correlation Maximization
- RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection
- GUDA: Counterfactual Group-wise Training Data Attribution for Diffusion Models via Unlearning
- COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics
- LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability
- Tagging-Augmented Generation: Assisting Language Models in Finding Intricate Knowledge In Long Contexts
- Omega-S: A Functional Resilience Index for LLM Fine-Tuning