MyanmarTTS

A from-scratch Burmese (Myanmar) text-to-speech model. 31M parameters, ~63 MB (fp16).

  • Language: Burmese (မြန်မာဘာသာ)
  • Training data: ~1.34M samples (news + real-world audio)
  • Training steps: 81,000
  • Training hardware: A100-80GB (Colab Pro+), ~13 hours, ~88 compute units
  • Inference hardware: Free Colab T4, any modern GPU, or CPU (slower)
  • Architecture: StableTTS (DiT + flow matching) + Vocos vocoder
  • License: CC0 1.0 (public domain, no attribution required)

Sample Output

All samples generated with default settings (euler solver, 12 steps).

Sample 0

Text: "မြန်မာလူမျိုးများဟာ အလွန် ယဥ်ကျေးသိမ်မွေ့ပြီး ဧည့်သည်များကို ပျူပျူငှာငှာ လှိုက်လှိုက်လှဲလှဲနဲ့ ကြိုဆိုကြပါတယ်"

Sample 1

Text: "မင်္ဂလာပါရှင် ကျွန်မကတော့ မြန်မာလူမျိုး ကရင်တိုင်းရင်းသူ အမျိုးသမီးလေး တစ်ဦး ဖြစ်ပါတယ်"

Sample 2

Text: "ဒီနေ့ ကျွန်မတို့ရဲ့ တီတီအက်စ် စနစ်သစ်လေး မော်ဒယ်အသစ်လေးတစ်ခုကို အောင်မြင်စွာ လေ့ကျင့် သင်ကြားနိုင်ခဲ့ပါတယ်"

Sample 3

Text: "လူသားတိုင်း လူသားတိုင်း ကိုယ်စိတ်နှစ်ဖြာ ကျန်းမာရွှင်လန်းပြီး စီးပွားလာဘ်လာဘတွေ ဒီရေအလား ကြီးပွား တိုးတက်နိုင်ကြပါစေ"

Sample 4

Text: "ဒီအသံထုတ်စနစ်လေးကို အသုံးပြုသူတိုင်း ကျန်းမာချမ်းသာပြီး လိုရာဆန္ဒတွေ တလုံးတဝတည်း ပြည့်စုံနိုင်ကြပါစေ"

Sample 5

Text: "ဟယ်လို... ဒါလင်... မတွေ့ရတာ... ကြာပီနော်"

Reference Audio

The reference voice used for voice cloning during inference:

Quick Start

import torch, soundfile as sf
from huggingface_hub import hf_hub_download
from api import StableTTSAPI
from burmese import burmese_to_ipa2

# Download model assets
model_path = hf_hub_download("freococo/MyanmarTTS", "model_fp16.pt")
vocos_path = hf_hub_download("freococo/MyanmarTTS", "vocos.pt")
ref_path   = hf_hub_download("freococo/MyanmarTTS", "samples/sample_0.wav")

# Initialize model
model = StableTTSAPI(model_path, vocos_path, "vocos").to("cuda")
model.g2p_mapping["burmese"] = burmese_to_ipa2

# Inference
text = "လူသားတွေ အားလုံးကို အရမ်း ချစ်ပါတယ်ရှင့်"
audio, _ = model.inference(text, ref_path, "burmese", step=32, solver="dopri5", cfg=3.0)
sf.write("output.wav", audio.squeeze(0).cpu().numpy(), 44100)

Translation: "I love all human beings very much." (female polite form)


Default Settings

  • Solver: euler (fast, ~16× faster than dopri5)
  • Steps: 12 (RTF ~0.08 on T4 GPU - 12× faster than real-time)
  • CFG: 3.0

Override per call:

audio = tts.tts("ဟယ်လို ဒါလင် မတွေ့ရတာ ကြာပီ", step=8)                     # fastest
audio = tts.tts("ဟယ်လို ဒါလင် မတွေ့ရတာ ကြာပီ", step=24)                    # higher quality
audio = tts.tts("ဟယ်လို ဒါလင် မတွေ့ရတာ ကြာပီ", solver="dopri5", step=32)   # max quality

Files

File Size Purpose License
model_fp16.pt 63 MB Default model weights CC0
model_fp32.pt 126 MB Full precision weights CC0
vocos.pt 57 MB Mel-to-waveform vocoder MIT (KdaiP)
vocab.txt 98 tokens Burmese character-level vocab CC0
config.json -- Mel + model configuration CC0
symbols.py, burmese.py -- Text frontend & G2P MIT (adapted)
api.py -- Inference wrapper API MIT (adapted)
samples/ -- Demo audio WAV files CC0
transcripts.json -- Sample text/audio mapping CC0
NOTICE -- Full license summary --

Training Progression

See checkpoint_comparison for sample audio generated at steps 8k, 33k, 44k, 55k, and 81k.


Acknowledgments

This work would not exist without the generous open-source community and AI assistance:

  • StableTTS by KdaiP — The DiT + flow-matching architecture and training code (MIT).
  • Vocos — Pretrained mel-to-wav vocoder (MIT).
  • DeepSeek AI — Provided AI pair-programming and engineering assistance throughout the project. From data pipeline design and architecture choices to resolving CUDA OOM bottlenecks, DeepSeek's guidance was instrumental at every stage.
  • The Burmese open-data community — For providing the audio corpora that made training possible.

Special Thanks

To DeepSeek AI — a true engineering partner from the first line of code to the final deployment. This model exists because of that collaboration.


License

  • Model weights and generated audio: CC0 1.0 Universal (Public Domain).
  • Supporting code: Adapted from KdaiP/StableTTS (MIT).
  • Vocoder: From KdaiP/StableTTS1.1 (MIT).

See NOTICE for additional details.

Downloads last month
113
Safetensors
Model size
31.6M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support