Skip to content
AI.info
News
Tools
Research
Learn
AI.info
Wanggui He
Explore Wanggui He on AI.info.
Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO
Page 1