view article Article Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original MultiverseComputingCAI • 5 days ago • 36
view article Article Tricks from OpenAI gpt-oss YOU 🫵 can use with transformers +5 ariG23498, sergiopaniego, reach-vb, pcuenq, ArthurZ, SaylorTwift, cyrilvallez • Sep 11, 2025 • 189
view article Article From Zero to GPU: A Guide to Building and Scaling Production-Ready CUDA Kernels drbh, danieldk • Aug 18, 2025 • 110
view article Article Accelerate ND-Parallel: A guide to Efficient Multi-GPU Training +3 smohammadi, siro1, winglian, marcsun13, djsaunde • Aug 8, 2025 • 100
MrezaPRZ/realmath_2025-2025-02-with-gemini-2_5-flash-context-2025-06-23-hard-cleaned Viewer • Updated Jul 25, 2025 • 105 • 12
MrezaPRZ/realmath_2025-2025-02-with-gemini-2_5-flash-context-2025-06-23-hard-cleaned Viewer • Updated Jul 25, 2025 • 105 • 12