AI.info
Zheng Li
Explore Zheng Li on AI.info.
- HeaPA: Difficulty-Aware Heap Sampling and On-Policy Query Augmentation for LLM Reinforcement Learning
- Compensating Distribution Drifts in Class-incremental Learning of Pre-trained Vision Transformers
- Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
- Instant Personalized Large Language Model Adaptation via Hypernetwork
- DeepPlanner: Scaling Planning Capability for Deep Research Agents via Advantage Shaping