We're going to release our BananaMind 2.1 models very soon! We're also announcing 2 new models.
All of our models we will train are: BananaMind 2.1 Flash Lite, 10M parameters with 8M in transformer and 2M in n-gram. 50B pretraining tokens. BananaMind 2.1 Lite with 25M parameters, 5M in n-gram and 20M in transformer. 75B pretraining tokens. BananaMind 2.1 Flash with 50M parameters, with undecided n-gram count yet. 100B pretraining tokens. BananaMind 2.1 Pro with 145M parameters, with undecided n-gram count yet. 150-200B pretraining tokens. BananaMind 2.1 Coder with 149M parameters with undecided n-gram count yet. We're now announcing BananaMind 2.1 NanoCoder, a 10M parameter model focused specifically on coding and BananaMind 2.1 MiniCoder which is a 25M parameter model focused on coding.
We're excited to release BananaMind 2.1 Pico Preview!
It includes the first preview of our BananaMind 2.1 architecture! This model gets near BananaMind 2 Micro performance at half the parameters and 37.5x less tokens! Thats insane!
The current architectures includes about 500K parameters of the total 1.5M parameters in n-gram embeddings and the layer 2 is run twice.
It also includes XSA and the XSA refresh gate.
We're still going to improve the architecture in the final release.
- Complete modern UI redesign - Adds support for Qwen3.5 0.8B, LFM2.5 230M,350M, SmolLM2 360M, Gemma 3 270M. - Adds Q7,Q6,Q5,Q3,Q1 quantization formats with a easy to use precision slider - And more!
The new UI includes: - New 1024ร768 High Quality interface. - Photographic QOI background. - Transparent BananaMind, CPU, cube, mouse, and Send icons. - Proper bitmap cursor. - Rounded translucent panels and cards. - Modern model-loading progress window. - Redesigned inference screen with response and prompt panels. - Localized redraws for the cursor, clicks, loading progress, and precision slider.
Notice: Qwen3.5 0.8B currently generates garbled text, it will be fixed tomorrow.
We're delaying BananaMind 2.1! When BananaMind 2.1 Lite was almost done, we benchmarked it and the results we're worse than BananaMind 2 Mini.
We're going to spend alot more time in research on tiny models and then scaling up our techniques to the actual BananaMind 2.1 models!
We're also announcing these new models: BananaMind 2.1 Coder: A 149M instruction tuned coder model trained on 75B tokens + 10B tokens of stack-v3-train. BananaMind 2.1 Pico: A 1M parameter model trained on 22B tokens of data. We also may release BananaMind 2.1 Large with around 100M parameters depending on how much compute we have.