shaheen6's picture
Publish Pixel Translation checkpoints
7a4e276
|
Raw History Blame Contribute Delete
1.81 kB
metadata
library_name: pytorch
pipeline_tag: translation
tags:
  - arabic
  - english
  - machine-translation

Pixel Translation checkpoints

Final checkpoints for the Base, Vocab+, Byte, and Contextual Pixel Arabic-to-English translation models. These weights are intended for use with mbzuai-nlp/arapixel-mt.

Layout

Each model has MSA-pretrained checkpoints for 5%, 10%, 25%, 50%, and the full MSA training set:

<model>/msa/{5pct,10pct,25pct,50pct,full}/pytorch_model.bin

The curriculum checkpoints were fine-tuned on the 1:1 MSA/dialect mixture:

<model>/finetuned/{5pct,10pct,25pct,50pct}/msa_dialect_1x/best_model.bin

Full-data checkpoints have four fine-tuning variants:

<model>/finetuned/full/dialect/best_model.bin
<model>/finetuned/full/msa_dialect_1x/best_model.bin
<model>/finetuned/full/msa_dialect_2x/best_model.bin
<model>/finetuned/full/msa_dialect_4x/best_model.bin

<model> is one of pixel, bytes, vocab, or base. In the codebase, the Base model is selected with --model decoder.

Every checkpoint directory also contains training_complete.json, which lets evaluation select the fixed MSA or dynamic fine-tuning source layout. The shared Vocab+ tokenizer is stored in vocab/tokenizer/.

Download one checkpoint

hf download shaheen6/PixelTranslationCheckpoints   --include "pixel/msa/full/*"   --local-dir checkpoints

The downloaded model path is then:

checkpoints/pixel/msa/full/pytorch_model.bin

Pass that file to --model-checkpoint. For Vocab+, also download the tokenizer:

hf download shaheen6/PixelTranslationCheckpoints   --include "vocab/tokenizer/*"   --local-dir checkpoints

export VOCAB_TOKENIZER_DIR="$PWD/checkpoints/vocab/tokenizer"