AI.info
Latest Research — Page 115
Browse Latest Research on AI.info.
- Spec-o3: A Tool-Augmented Vision-Language Agent for Rare Celestial Object Candidate Vetting via Automated Spectral Inspection
- FACE: Faithful Automatic Concept Extraction
- StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
- From Euclidean to Graph-Structured Data: A Survey of Collaborative Learning
- Cross Modal Fine-Grained Alignment via Granularity-Aware and Region-Uncertain Modeling
- ImprovEvolve: Ask AlphaEvolve to Improve the Input Solution and Then Improvise
- Scaling Test-Time Compute to Achieve IOI Gold Medal with Open-Weight Models
- Condition Matters in Full-head 3D GANs
- StudentBench: AI and human tutoring yield equivalent GRE learning gains
- Operator-Based Generalization Bound for Deep Learning: Insights on Multi-Task Learning
- Comparison requires valid measurement: Rethinking attack success rate comparisons in AI red teaming
- FS-Researcher: Test-Time Scaling for Long-Horizon Research Tasks with File-System-Based Agents
- Provable Memory Efficient Self-Play Algorithm for Model-free Reinforcement Learning
- How Well Do LLMs Understand Drug Mechanisms? A Knowledge + Reasoning Evaluation Dataset
- ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering
- WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenarios
- Fun-Audio-Chat Technical Report
- Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks
- CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport
- CABLE: Extending the Reach of Memory Retrieval via Complementary Antecedent-Based Linking and Expansion
- UIKA: Fast Universal Head Avatar from Pose-Free Images
- Balanced Learning for Domain Adaptive Semantic Segmentation
- SHRUG-FM: Reliability-Aware Foundation Models for Earth Observation
- DiaVLo: Diagnosing Behaviours of Vision-Language Models
- SenTSR-Bench: Thinking with Injected Knowledge for Time-Series Reasoning
- MARS: Modular Agent with Reflective Search for Automated AI Research
- Privacy-enhanced federated learning via asynchronous aggregation and local differential perturbation
- MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning
- DP$^2$O-SR: Direct Perceptual Preference Optimization for Real-World Image Super-Resolution
- LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization
- SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning
- RFEval: Benchmarking Reasoning Faithfulness under Counterfactual Reasoning Intervention in Large Reasoning Models
- FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation
- ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment
- Breaking the Gradient Barrier: Unveiling Large Language Models for Strategic Classification
- PubSub-VFL: Towards Efficient Two-Party Split Learning in Heterogeneous Environments via Publisher/Subscriber Architecture
- Optimizing Diversity and Quality through Base-Aligned Model Collaboration
- Evaluating LLM Story Generation through Large-scale Network Analysis of Social Structures
- Don't Adapt Small Language Models for Tools; Adapt Tool Schemas to the Models
- MPJudge: Towards Perceptual Assessment of Music-Induced Paintings