Ill wait for the glm distilled xD
I think we are years away though
Ill wait for the glm distilled xD
I think we are years away though
Wild if true
Feeling the same :)
Imagine having some kind of server so it always feels updated with your latest integrated models!
byte-level tokenizer is the part that caught my attention, keep it up sir!
Let's go!!!
Yes sir!
Do you mean like the architecture? Could you point at the model?
Waiting for your new models sir
Hey @appvoid
Are you gonna press it?
Done. hahaha bots are becoming more of a thing here lately.
Hmm interesting!
You should only use some loops, like lets say you have 4 layers then:
L1 🠮 L2 🠮 L2 🠮 L3 🠮 L4 so 5 effective layers notL1 🠮 L1 🠮 L2 🠮 L2 🠮 L3 🠮 L3 🠮 L4 🠮 L4 because our tests on 1M:
Metric All looped (6 blocks) Partial looped (4 blocks) No loop (3 blocks) Base Bench Elo 875 885 884 Base raw accuracy 33.14% 34.57% 33.71% Base weighted accuracy 32.54% 33.65% 33.57% ARC Easy acc_norm 30.98% 29.50% 30.35% ARC Challenge acc_norm 21.33% 22.53% 22.10% PIQA acc_norm 52.45% 54.30% 52.56% HellaSwag acc_norm 27.28% 27.04% 26.93% ArithMark-3 acc_norm 30.40% 33.00% 30.80% INT Index 3.88 5.37 3.93 Training throughput 344K tok/s 492K tok/s 553K tok/s
It's always 4 steps, every time. I don't know why but I think it has something to do with model capacity.
Saying "Nobody knows what they are doing" is just a convenient excuse to justify terrible engineering. There is a fine line between scientific trial-and-error and proud, brute-force ignorance.Let's be clear about your "frontier":Blind Gambling: When independent labs don't understand the underlying mathematics or hardware physics, they just throw data at a wall and pray to the loss curve. That is digital alchemy, not science.The Loop: Instead of fixing structural bottlenecks or learning non-linear dynamics, people just brute-force configs. It’s the engineering equivalent of a cat grooming itself because it has nothing else to do.Zero Legacy: This unscientific approach is why the ecosystem is flooded with overfitted, hollow checkpoints that break down outside of their strict test sets.You aren't advancing the frontier; you are just polluting the platform because you refuse to open a textbook. Brute force has hit a physical wall. True innovation requires cognitive architecture and actual engineering, not just romanticizing failure. 🫵🤡
Lol. AI is being used by the very trolls that try to destroy it. Quite ironic. Quality bait though.
Basically, they nerfed their models, then they lower limits without any transparency. Finally, they accused their users of being using their models the wrong way. Gave their free users luna by default and nerfed back to 5-hours limit plus users.