phoonnx-orpheus β€” Orpheus 3B (English), ONNX

ONNX weights for Orpheus TTS by Canopy Labs, packaged for phoonnx. This repository holds converted weights only β€” no new training was done. Original model and code are Apache-2.0; all credit for the model belongs to Canopy Labs.

Orpheus is a Llama-3.2-3B causal LM whose vocabulary carries 28 672 audio tokens. It emits a flat token stream; every seven tokens form one SNAC frame, which SNAC's decoder turns into 2048 samples at 24 kHz.

Read this before you use it: Orpheus is not a CPU model

A 3B backbone needs about 0.37 s per decode step on 12 CPU cores, and SNAC needs ~82 tokens for every second of audio. That is roughly 37-41x slower than real time on CPU. Canopy Labs announced 1B / 400M / 150M tiers but never released them β€” their own loader still raises "not supported ... will be released very soon" β€” so there is no smaller Orpheus to fall back to.

Use this on a GPU. It is mirrored for completeness, not as an offline CPU default.

Files (orpheus-3b-en-onnx/)

File What it is
model.onnx + model.onnx_data..model.onnx_data_6 Llama backbone, fp32, KV-cached. The graph needs all seven .onnx_data shards next to it.
snac_decoder.onnx SNAC 24 kHz decoder, fp32, copied from onnx-community/snac_24khz-ONNX (Apache-2.0), unmodified.
tokenizer.json The checkpoint's own BPE, copied from upstream.

model.onnx comes from onnx-community/orpheus-3b-0.1-ft-ONNX, re-verified against the torch reference before mirroring (see below).

Voices

tara, leah, jess, leo, dan, mia, zac, zoe β€” in Canopy Labs' own order of conversational realism. The voice is a name written into the prompt text, not an embedding or a speaker id.

Emotive tags are ordinary text the same BPE encodes: <laugh>, <chuckle>, <sigh>, <cough>, <sniffle>, <groan>, <yawn>, <gasp>.

Verification

Measured against the torch reference (unsloth/orpheus-3b-0.1-ft, an ungated copy of the gated original) on a 19-token served prompt, fp32 both sides:

Variant Size Prefill max abs logit diff Greedy agreement, 25 steps ms / decode step Verdict
model (fp32) 12.7 GB 0.166 25/25 369 mirrored
model_q4 2.4 GB 8.65 23/25 371 rejected β€” loses parity and is not faster
model_q4f16 2.1 GB 8.53 10/25 96 rejected β€” fails parity

Only fp32 is mirrored. The quantized variants either lose the greedy path or, in the q4 case, buy nothing: its weights are int4 but its activations stay fp32, so onnxruntime dequantizes on the fly and the step cost is unchanged.

SNAC's decoder is stochastic β€” it contains a noise block, so two decodes of the same codes differ. Its ONNX export was therefore checked against the model's own run-to-run spread (mean of 40 decodes, relative RMS) rather than by exact diff:

SNAC decoder Relative RMS vs torch mean Ratio to noise floor Verdict
decoder_model (fp32) 0.0200 0.89x pass
decoder_model_fp16 0.0233 1.03x pass
decoder_model_int8 0.1356 6.02x fail
decoder_model_uint8 / _quantized 0.1089 4.83x fail

decoder_model_q4 and decoder_model_bnb4 are byte-for-byte the same size as fp32 and produce identical output β€” the quantizer skipped this convolutional graph entirely, so those names are misleading rather than smaller.

Licensing

  • Orpheus model and code: Apache-2.0, Β© Canopy Labs.
  • SNAC: Apache-2.0, Β© Hubert Siuzdak.
  • This repository: Apache-2.0. Converted weights only.

Canopy Labs ask that the model not be used for impersonation without consent, misinformation, or deception.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for OpenVoiceOS/phoonnx-orpheus

Collection including OpenVoiceOS/phoonnx-orpheus