Pixel Translation checkpoints
Final checkpoints for the Base, Vocab+, Byte, and Contextual Pixel Arabic-to-English translation models. These weights are intended for use with mbzuai-nlp/arapixel-mt.
Layout
Each model has MSA-pretrained checkpoints for 5%, 10%, 25%, 50%, and the full MSA training set:
<model>/msa/{5pct,10pct,25pct,50pct,full}/pytorch_model.bin
The curriculum checkpoints were fine-tuned on the 1:1 MSA/dialect mixture:
<model>/finetuned/{5pct,10pct,25pct,50pct}/msa_dialect_1x/best_model.bin
Full-data checkpoints have four fine-tuning variants:
<model>/finetuned/full/dialect/best_model.bin
<model>/finetuned/full/msa_dialect_1x/best_model.bin
<model>/finetuned/full/msa_dialect_2x/best_model.bin
<model>/finetuned/full/msa_dialect_4x/best_model.bin
<model> is one of pixel, bytes, vocab, or base. In the codebase,
the Base model is selected with --model decoder.
Every checkpoint directory also contains training_complete.json, which lets
evaluation select the fixed MSA or dynamic fine-tuning source layout. The shared
Vocab+ tokenizer is stored in vocab/tokenizer/.
Download one checkpoint
hf download shaheen6/PixelTranslationCheckpoints --include "pixel/msa/full/*" --local-dir checkpoints
The downloaded model path is then:
checkpoints/pixel/msa/full/pytorch_model.bin
Pass that file to --model-checkpoint. For Vocab+, also download the tokenizer:
hf download shaheen6/PixelTranslationCheckpoints --include "vocab/tokenizer/*" --local-dir checkpoints
export VOCAB_TOKENIZER_DIR="$PWD/checkpoints/vocab/tokenizer"