AI.info
Latest Research — Page 34
Browse Latest Research on AI.info.
- RGMP: Recurrent Geometric-prior Multimodal Policy for Generalizable Humanoid Robot Manipulation
- Pb4U-GNet: Resolution-Adaptive Garment Simulation via Propagation-before-Update Graph Network
- Image Generation as a Visual Planner for Robotic Manipulation
- Rethinking Deep Alignment Through The Lens Of Incomplete Learning
- DC4GS: Directional Consistency-Driven Adaptive Density Control for 3D Gaussian Splatting
- WaiT for the Signal: Simple Frequency-Aware Flow-Matching
- Enhancing Adversarial Transferability by Balancing Exploration and Exploitation with Gradient-Guided Sampling
- RadarGen: Automotive Radar Point Cloud Generation from Cameras
- Compositional Steering of Large Language Models with Steering Tokens
- MM-Spectrum: Multimodal Multi-spectral Molecular Structural Elucidation with a Stable MoE Framework
- Peek2: Regex-free Byte-level Byte-Pair Encoding Pretokenizer for LLM Inference on Edge Devices
- XLinear: A Lightweight and Accurate MLP-Based Model for Long-Term Time Series Forecasting with Exogenous Inputs
- Complexity as Advantage: A Regret-Based Perspective on Emergent Structure
- The Deleuzian Representation Hypothesis
- Text to Sketch Generation with Multi-Styles
- An Epistemic Perspective on Agent Awareness
- AccuQuant: Simulating Multiple Denoising Steps for Quantizing Diffusion Models
- Cross-Domain Few-Shot Segmentation via Multi-view Progressive Adaptation
- Self-Refining Video Sampling
- Ovis-Image Technical Report
- ABBSPO: Adaptive Bounding Box Scaling and Symmetric Prior based Orientation Prediction for Detecting Aerial Image Objects
- SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
- Depth-Supervised Fusion Network for Seamless-Free Image Stitching
- SARMAE: Masked Autoencoder for SAR Representation Learning
- Resurfacing Paralinguistic Awareness in Large Audio Language Models
- UniMedVL: Unifying Medical Multimodal Understanding and Generation through Observation-Knowledge-Analysis
- Re$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning
- AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
- R-Align: Enhancing Generative Reward Models through Rationale-Centric Meta-Judging
- Gated Tree Cross-Attention for Checkpoint-Compatible Syntax Injection in Decoder-Only LLMs
- Beyond Monotonicity: Revisiting Factorization Principles in Multi-Agent Q-Learning
- Text-Conditioned Background Generation for Editable Multi-Layer Documents
- Ontology-supported AI Model and Dataset Management
- DiffMM: Efficient Method for Accurate Noisy and Sparse Trajectory Map Matching via One Step Diffusion
- Think Twice: Branch-and-Rethink Reasoning Reward Model
- MoETTA: Test-Time Adaptation Under Mixed Distribution Shifts with MoE-LayerNorm
- WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigation
- Length-MAX Tokenizer for Language Models
- Vision-Language Memory for Spatial Reasoning
- Beyond the Lower Bound: Bridging Regret Minimization and Best Arm Identification in Lexicographic Bandits