Matched OpenR1 SFT checkpoints and subsequent GRPO models for studying token-level behavior during the transition from SFT to RL.
🏝️ On Vacation
Zbib M
zbeeb
AI & ML interests
KAUST - AUB
Recent Activity
upvoted a paper about 2 hours ago
Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation updated a collection about 3 hours ago
OpenR1 SFT to GRPO: Token-Level Study updated a model about 3 hours ago
zbeeb/Qwen2.5-Math-1.5B-OpenR1-SFT-GRPOOrganizations
Staleness: GRPO Models & Data
GRPO checkpoints organized by staleness cap, with the shared 17,005-row DAPO math dataset, provenance and benchmark scores.
-
zbeeb/Qwen2.5-Math-1.5B-GRPO-Staleness-2
Text Generation • 2B • Updated • 508 -
zbeeb/Qwen2.5-3B-GRPO-Staleness-2
Text Generation • 3B • Updated • 526 -
zbeeb/Qwen2.5-Math-7B-GRPO-Staleness-2
Text Generation • 8B • Updated • 531 -
zbeeb/Qwen2.5-Math-1.5B-GRPO-Staleness-4
Text Generation • 2B • Updated • 194
Reasoning Vectors
OpenR1 SFT to GRPO: Token-Level Study
Matched OpenR1 SFT checkpoints and subsequent GRPO models for studying token-level behavior during the transition from SFT to RL.
Easy-RL: Verifiable and cleaned Math Data
Deduplicated Skywork and DeepScaleR math pools, plus their merged version. Only Math-Verify parser-compatible references.
Staleness: GRPO Models & Data
GRPO checkpoints organized by staleness cap, with the shared 17,005-row DAPO math dataset, provenance and benchmark scores.
-
zbeeb/Qwen2.5-Math-1.5B-GRPO-Staleness-2
Text Generation • 2B • Updated • 508 -
zbeeb/Qwen2.5-3B-GRPO-Staleness-2
Text Generation • 3B • Updated • 526 -
zbeeb/Qwen2.5-Math-7B-GRPO-Staleness-2
Text Generation • 8B • Updated • 531 -
zbeeb/Qwen2.5-Math-1.5B-GRPO-Staleness-4
Text Generation • 2B • Updated • 194
TAPS
Reasoning Vectors