Prompt Cache: Modular Attention Reuse for Low-Latency Inference Paper • 2311.04934 • Published Nov 7, 2023 • 33
Running on CPU Upgrade Agents Featured 1.01k Model Memory Utility 🚀 1.01k Calculate GPU memory needed for training Hugging Face models
Running on CPU Upgrade 14.1k Open LLM Leaderboard 🏆 14.1k Track, rank and evaluate open LLMs and chatbots