AI.info
Yangqiu Song
Explore Yangqiu Song on AI.info.
- AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon Conversations
- HeaPA: Difficulty-Aware Heap Sampling and On-Policy Query Augmentation for LLM Reinforcement Learning
- $\mathbb{R}^{2k}$ is Theoretically Large Enough for Embedding-based Top-$k$ Retrieval
- WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality
- AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
- DeepPlanner: Scaling Planning Capability for Deep Research Agents via Advantage Shaping
- DixitWorld: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay
- NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents