Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published Sep 3 • 104
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published Aug 31 • 97
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving Paper • 2609.00111 • Published Aug 31 • 316
PaperBanana: Automating Academic Illustration for AI Scientists Paper • 2601.23265 • Published Jan 30 • 230
LaViPlan : Language-Guided Visual Path Planning with RLVR Paper • 2507.12911 • Published Jul 17, 2025
QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs Paper • 2510.11696 • Published Oct 13, 2025 • 183
A Survey of Reinforcement Learning for Large Reasoning Models Paper • 2509.08827 • Published Sep 10, 2025 • 193