TogetherAI 开源 OSCAR:超越 TurboQuant! 面向真实 Serving 的 2-bit KV Cache 量化
中文摘要
Together AI开源OSCAR量化方案,实现2-bit KV Cache,显著降低长文本推理显存并提升吞吐量,同时保持与BF16相当的精度。
English Summary
Together AI open-sourced OSCAR, a 2-bit KV Cache quantization method that reduces memory usage and boosts throughput for long-context inference while maintaining BF16-level performance.
原文节选
📌 一句话摘要 Together AI 开源了 OSCAR,一种面向真实长上下文服务的 2-bit KV Cache 量化方案,通过注意力感知的旋转技术,在显著降低显存占用和提升推理吞吐的同时,保持了与 BF16 精度相当的模型性能。 📝 详细摘要 本文详细介绍了 Together AI 开源的 OSCAR(Offline Spectral Covariance-Aware Rotation)方案,旨在解决长上下文 L...