Matched OpenR1 SFT checkpoints and subsequent GRPO models for studying token-level behavior during the transition from SFT to RL.
🏝️ On Vacation
Zbib M
zbeeb
AI & ML interests
KAUST - AUB
Recent Activity
updated a collection about 5 hours ago
OpenR1 SFT to GRPO: Token-Level Study updated a collection about 5 hours ago
OpenR1 SFT to GRPO: Token-Level Study updated a dataset about 5 hours ago
zbeeb/OpenR1-SFT-Math-20kOrganizations
OpenR1 SFT to GRPO: Token-Level Study
Matched OpenR1 SFT checkpoints and subsequent GRPO models for studying token-level behavior during the transition from SFT to RL.
Easy-RL: Verifiable and cleaned Math Data
Deduplicated Skywork and DeepScaleR math pools, plus their merged version. Only Math-Verify parser-compatible references.
models 30
zbeeb/Qwen2.5-3B-OpenR1-SFT
Text Generation • 3B • Updated
zbeeb/Qwen2.5-Math-1.5B-OpenR1-SFT
Text Generation • 2B • Updated
zbeeb/Qwen2.5-1.5B-OpenR1-SFT
Text Generation • 2B • Updated
zbeeb/Qwen2.5-Math-7B-GRPO-Staleness-4
Text Generation • 8B • Updated • 122
zbeeb/Qwen2.5-3B-GRPO-Staleness-6
Text Generation • 3B • Updated • 226
zbeeb/Qwen2.5-3B-GRPO-Staleness-8
Text Generation • 3B • Updated • 197
zbeeb/Qwen2.5-Math-1.5B-GRPO-Staleness-8
Text Generation • 2B • Updated • 302
zbeeb/Qwen2.5-3B-GRPO-Staleness-4
Text Generation • 3B • Updated • 378
zbeeb/Qwen2.5-Math-1.5B-GRPO-Staleness-6
Text Generation • 2B • Updated • 420
zbeeb/Qwen2.5-Math-1.5B-GRPO-Staleness-4
Text Generation • 2B • Updated • 191
datasets 13
zbeeb/OpenR1-SFT-Math-20k
Viewer • Updated • 60.8k
zbeeb/Skywork-DeepScaleR-Merged-Verifiable-Dedup
Viewer • Updated • 98.9k • 40
zbeeb/DeepScaleR-Verifiable-Dedup
Viewer • Updated • 37.7k • 41
zbeeb/Skywork-OR1-Math-Verifiable-Dedup
Viewer • Updated • 97.8k • 48
zbeeb/Staleness-GRPO-DAPO-Math-17k
Viewer • Updated • 17k • 133
zbeeb/dapo-math-17k-qwen3-clean
Viewer • Updated • 17k • 29
zbeeb/beneficial-dpo-dataset
Viewer • Updated • 10k • 36
zbeeb/TAPS-Datasets
Viewer • Updated • 210k • 41 • 4
zbeeb/mixeddatasafety7030
Viewer • Updated • 10k • 2
zbeeb/llama3_MathInstruct_data
Updated • 6