AI.info
Latest Research — Page 117
Browse Latest Research on AI.info.
- Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents
- VL-RouterBench: A Benchmark for Vision-Language Model Routing
- Calibrating Uncertainty for Zero-Shot Adversarial CLIP
- Evaluating Online Moderation Via LLM-Powered Counterfactual Simulations
- SIGN: Schema-Induced Games for Naming
- Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification
- Fewer Weights, More Problems: A Practical Attack on LLM Pruning
- TSVer: A Benchmark for Fact Verification Against Time-Series Evidence
- Unveiling the Spatial-temporal Effective Receptive Fields of Spiking Neural Networks
- SENSE: Self-Supervised Neural Embeddings for Spatial Ensembles
- LOOM: Personalized Learning Informed by Daily LLM Conversations Toward Long-Term Mastery via a Dynamic Learner Memory Graph
- PHyCLIP: $\ell_1$-Product of Hyperbolic Factors Unifies Hierarchy and Compositionality in Vision-Language Representation Learning
- Toward Inclusive Avatar Design with Limb Differences Through Artificial Intelligence
- Post-Training LLMs as Better Decision-Making Agents: A Regret-Minimization Approach
- Enhancing Object Detection with Privileged Information: A Model-Agnostic Teacher-Student Approach
- MUG-V 10B: High-efficiency Training Pipeline for Large Video Generation Models
- Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers
- Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment Formats
- Native Hybrid Attention for Efficient Sequence Modeling
- A Simple Method to Enhance Pre-trained Language Models with Speech Tokens for Classification
- A Probabilistic U-Net Approach to Downscaling Climate Simulations
- NodeImport: Imbalanced Node Classification with Node Importance Assessment
- It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief
- TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning Tasks
- D-GARA: A Dynamic Benchmarking Framework for GUI Agent Robustness in Real-World Anomalies
- Co-design of Neural and Muscle Network based on Embodied Perceptron Representation
- Zero-Shot Transfer Capabilities of the Sundial Foundation Model for Leaf Area Index Forecasting
- Demo: Generative AI helps Radiotherapy Planning with User Preference
- Transferable Black-Box One-Shot Forging of Watermarks via Image Preference Models
- CLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space
- SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal Bias
- SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios
- Pointy - A Lightweight Transformer for Point Cloud Foundation Models
- GEOID-Flood: A Large-Scale Multi-Modal Benchmark Dataset for Flood Segmentation
- Probing Multimodal Large Language Models on Cognitive Biases in Chinese Short-Video Misinformation
- G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
- A Fully First-Order Layer for Differentiable Optimization
- ReasonEdit: Towards Reasoning-Enhanced Image Editing Models
- Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum
- Revisiting the UID Hypothesis in LLM Reasoning Traces