"In the beginning was the word. Then came the **** word processor. Then came the thought processor. Then came the death of literature. And so it goes" -- Hyperion - Dan Simmons
Today, we are announcing a brand-new series of SupraLabs models: Supra2 This series will feature various models, including such as: - ๐ Supra2-Nano (0.4M) โ The smallest Supra2 model. - ๐ค Supra2-Small (1.4M) โ The tiny model that runs everywhere. - ๐ช Supra2-Medium (25M) โ Our medium class model in the Supra2 family. The powerful midsizer. - ๐ฅ Supra2-Pro (100M): base, instruct, reasoning, code, math and more! โ The most capable model yet! A real allrounder for all your everyday tasks. - ๐จ Supra2-IMG โ our generative text-to-image model ...and many more...
Current progress: - Nano (0.4M) and Small (1.4M): in training; almost done. Baseline set. - Medium (25M): coming soon... - Pro (100M): in training; finishes in 66 hours - Monday, 3rd August 2026, 12:00AM - IMG: coming soon...
You can support us with a like and follow if you want! Don't miss our next release! Stay tuned...
Introducing Unsloth for AMD ๐ You can now train & run LLMs on your AMD hardware
โข We collaborated with AMD to enable you to train & run 500+ models on AMD GPUs โข Works on Windows, WSL, Linux โข Train Qwen, Gemma on just 3GB VRAM
We wrote a full technical guide on how to train a bilingual (ES/EN) LLM from scratch: TinyQwen.
Covers: - Hybrid architecture based on Qwen3.5 - Pre-training with 15B tokens - Cost benchmark between H200 and B200 - Post-training with SFT + LoRA - Full code and data, open source
With ~$11 of compute on an H200 we ran an initial training run, enough to validate the full architecture and pipeline.