BlueCodec β€” speech autoencoder (codec only)

This repository publishes only the neural audio codec used by BlueTTS: a 44.1 kHz speech autoencoder that maps waveforms to a low-rate continuous latent sequence and back. It is not a full TTS model (no text encoder, duration model, or flow stack).

If you need… Use
End-to-end ONNX TTS notmax123/blue-onnx + BlueTTS
Full PyTorch stack + stats (training / voice export) notmax123/blue β€” includes blue_codec.safetensors alongside TTL/DP weights
Training the codec from scratch maxmelichov/blue-codec (standalone repo & training doc)

Project home: https://github.com/maxmelichov/BlueTTS Β· Live demo: Hugging Face Space β€” notmax123/Blue

What it does

  • Encoder: waveform β†’ spectrogram features β†’ 24-dimensional latents at ~86 Hz (compact trajectory for downstream TTS).
  • Decoder: latents β†’ high-quality 44.1 kHz audio (causal stack + vocoder head).

Downstream BlueTTS modules (flow matching, duration, text-to-latent) run in this latent space; keeping synthesis lightweight and fast.

Architecture (summary)

Piece Details
Input 1253-channel spectrogram (1025 log-linear + 228 log-mel; FFT 2048, hop 512)
Encoder (~25.6M params) Conv1d stem (1253β†’512) + 10 ConvNeXt blocks + projection (512β†’24)
Decoder (~25.3M params) CausalConv1d stem (24β†’512) + 10 causal dilated ConvNeXt blocks + vocoder head
Latent 24-D @ ~86 Hz

Checkpoint in this repo

File Role
model.safetensors BlueCodec 1.5M: encoder + decoder weights (Safetensors). State dict keys are typically prefixed with encoder.* and decoder.*.
encoder_supertonic3_decoder/encoder.safetensors Supertonic-3-decoder encoder (Sept 2026): encoder only (encoder.*), trained against the frozen official Supertonic-3 vocoder. The decoder is not included β€” see below.
encoder_supertonic3_decoder/ae_290000.pt Its training checkpoint (step 290k) without the decoder: encoder, discriminators, optimizers, schedulers.
encoder_supertonic3_decoder_edge_fixed/ Reserved for the edge-fixed encoder, a retrain that removes the end-of-clip edge code (in progress, not released).

*(An older naming convention in some local scripts is ae_latest.safetensors; the file served from this Hub repo is model.safetensors.)*

Encoder for the official Supertonic-3 vocoder (Sept 2026)

This is our encoder (initialised from the 1.5M-step encoder) trained for 290k steps against the frozen official Supertonic-3 vocoder, so its latents live in the official Supertonic-3 latent space. The vocoder weights are Supertone Inc.'s (Supertone/supertonic-3, BigScience OpenRAIL-M) and are not redistributed here; the bluecodec package downloads them from the official repo at load time:

from bluecodec import BlueCodec   # pip install onnx "bluecodec @ git+https://github.com/maxmelichov/blue-codec.git"
codec = BlueCodec.from_pretrained("notmax123/blue-codec", decoder="supertonic3")
z = codec.encode(audio, edge_pad_chunks=2)   # edge-padded encoding (see below)
y = codec.decode(z)[..., :audio.shape[-1]]

Round-trip audit on 10 reference clips (the project's historical log-mel L1 / 12–22 kHz energy vs source / DNSMOS): BlueCodec 1.5M 0.4426 / +19.56 dB / 3.40; this encoder + the official vocoder 0.3385 / +7.63 dB / 3.42.

Encode with edge padding. Bare encoding leaves an "edge code" in the last ~3 latent frames of a clip (the last compressed frame ~9x the median norm); it does not change the reconstruction, but downstream models copy it. encode(audio, edge_pad_chunks=2) removes it. Full analysis: technical report.

License. This encoder contains none of Supertone's weights, but was trained through their model and may count as a derivative under OpenRAIL-M; its use is subject to that license's use-based restrictions (a copy is in encoder_supertonic3_decoder/LICENSE.OpenRAIL-M).

Download

hf download notmax123/blue-codec --repo-type model --local-dir ./blue_codec_only

Equivalent:

huggingface-cli download notmax123/blue-codec --repo-type model --local-dir ./blue_codec_only

Repo id is case-sensitive: notmax123/blue-codec.

License

model.safetensors: MIT. encoder_supertonic3_decoder/: see the Supertonic-3 section above (OpenRAIL-M use-based restrictions apply). General:

MIT β€” align usage with BlueTTS and the blue-codec repository for any training or redistribution terms that apply to your use case.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
50.9M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Spaces using notmax123/blue-codec 2