AI.info
Latest Research — Page 76
Browse Latest Research on AI.info.
- EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
- STEP: Score-Based Temporal Energy for Human Pose Video Anomaly Detection
- SafeRedir: Prompt Embedding Redirection for Robust Unlearning in Image Generation Models
- Learning Joint Embeddings of Function and Process Call Graphs for Malware Detection
- ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image Retrieval
- Layer-Order Inversion: Rethinking Latent Multi-Hop Reasoning in Large Language Models
- Supervise Less, See More: Training-free Nuclear Instance Segmentation with Prototype-Guided Prompting
- Show-Harness: Just a VLM Agent Can Play Robots
- Lévy Attention: Single-Pass Predictive Uncertainty for Continuous-Time Attention
- D$^{3}$ToM: Decider-Guided Dynamic Token Merging for Accelerating Diffusion MLLMs
- PEDESTRIAN: An Egocentric Vision Dataset for Obstacle Detection on Pavements
- Unified Privacy Guarantees for Decentralized Learning via Matrix Factorization
- Reasoning for Social Audio-Visual Question Answering: Where Do We Stand?
- Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
- Rank-Learner: Orthogonal Ranking of Treatment Effects
- OFA-MAS: One-for-All Multi-Agent System Topology Design based on Mixture-of-Experts Graph Generative Models
- Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-Rollout
- Evaluating Hydro-Science and Engineering Knowledge of Large Language Models
- SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action Recognition
- Instance-Level Generation for Representation Learning
- Distributed Lag Neural Additive Models
- Connecting Jensen-Shannon and Kullback-Leibler Divergences: A New Bound for Representation Learning
- WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning
- Semi-Supervised Cross-Domain Imitation Learning
- Decoupled Entropy Minimization
- On Uncertainty Calibration for Equivariant Functions
- REM: Evaluating LLM Embodied Spatial Reasoning through Multi-Frame Trajectories
- EIDSeg: A Pixel-Level Semantic Segmentation Dataset for Post-Earthquake Damage Assessment from Social Media Images
- Vegas: Self-Speculative Decoding with Verification-Guided Sparse Attention
- MKSNet: Advanced Small Object Detection in Remote Sensing Imagery with Multi-Kernel and Dual Attention Mechanisms
- Breaking Scale Anchoring: Frequency Representation Learning for Accurate High-Resolution Inference from Low-Resolution Training
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- Sequential Multi-Agent Dynamic Algorithm Configuration
- IF:CARGO: LLM-Based Semantic Compilation for Al-Native Rule Programming Games
- CC-OR-Net: A Unified Framework for LTV Prediction through Structural Decoupling
- Generating the Past, Present and Future from a Motion-Blurred Image
- Learning to Factorize and Adapt: A Versatile Approach Toward Universal Spatio-Temporal Foundation Models
- IFlowNets: Extending Generative Samplers to Learn Strategies in Incomplete Information Games
- Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
- Revisiting the Uniform Information Density Hypothesis in LLM Reasoning