Owen Reed
owenreed
AI & ML interests
Efficient LLM inference, KV cache optimization, quantization, speculative decoding, model pruning
Recent Activity
upvoted a paper about 2 hours ago
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning upvoted a paper about 2 hours ago
Rethinking On-Policy Distillation of Large Language Models II: One Training Example liked a dataset about 20 hours ago
ergt2025/kvcache_offloadOrganizations
None yet