This model was fine‑tuned with GRPO for only 50 steps using 4 samples per step. The result is exceptionally high accuracy on JEE‑level mathematics problems, though its broader context handling and instruction‑following abilities were diminished. In essence, it has become a compact powerhouse — a “mini‑tank” built for raw mathematical problem‑solving rather than nuanced reasoning.

Downloads last month: 6

Safetensors

Model size

1.0B params

Tensor type

F16

Model tree for Parveshiiii/M1-MathX

Base model

google/gemma-3-1b-pt

Finetuned

google/gemma-3-1b-it

Finetuned

(471)

this model

Parveshiiii
/

M1-MathX

Model tree for Parveshiiii/M1-MathX

Dataset used to train Parveshiiii/M1-MathX