LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure Paper • 2608.13545 • Published 5 days ago • 5
SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning Paper • 2608.14277 • Published 4 days ago • 20
DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data Paper • 2608.13517 • Published 5 days ago • 19
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning Paper • 2608.14290 • Published 4 days ago • 30
Verifier-Induced Support Reshaping in On-Policy Optimization Paper • 2608.00220 • Published 18 days ago • 4
Scaling Domain Data Repetition in LLM Pretraining Paper • 2608.14071 • Published 4 days ago • 5
ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization Paper • 2608.11045 • Published 7 days ago • 3
Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing Paper • 2608.11660 • Published 6 days ago • 11
Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning Paper • 2607.29211 • Published 18 days ago • 12
SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models Paper • 2608.10538 • Published 7 days ago • 14
Parameter Exploration for RLVR via Variational Learning Paper • 2608.09805 • Published 7 days ago • 6
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published 8 days ago • 332
Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation Paper • 2608.10812 • Published 7 days ago • 14
On-Policy Self-Distillation without Any Supervision Paper • 2608.06296 • Published 9 days ago • 208
Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss Paper • 2608.03796 • Published 14 days ago • 14
OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching Paper • 2608.08097 • Published 10 days ago • 25