Post
52
We have released BGA!
And wow, It provides 256x (and 512x at the end of 1M) yes 256x LESS attention compute at 1M context window.
That means you can train a 1M context window at the compute of a ~4K context window.
Check IT OUT: BananaMind/blog
The Accuracy Should BE WAy better than DSA but untested yet.
And, now some updates on BananaMind 3:
BananaMind 3 Will start training Soon!
Sizes: 10M, 25M, 50M, 100M, 150M
And the context windows ARE INSANE: 10M, 16K context, 25M 16k context, 50M 32K context, 100M and 150M, 64K context!!!!
And wow, It provides 256x (and 512x at the end of 1M) yes 256x LESS attention compute at 1M context window.
That means you can train a 1M context window at the compute of a ~4K context window.
Check IT OUT: BananaMind/blog
The Accuracy Should BE WAy better than DSA but untested yet.
And, now some updates on BananaMind 3:
BananaMind 3 Will start training Soon!
Sizes: 10M, 25M, 50M, 100M, 150M
And the context windows ARE INSANE: 10M, 16K context, 25M 16k context, 50M 32K context, 100M and 150M, 64K context!!!!