The quant that shrank my 27B model by a third — GSQ-RCO, explained without the tensor algebra
中文摘要
GSQ-RCO技术可将27B模型缩小30%,在优化比特分配的同时保持模型质量。
English Summary
GSQ-RCO quantization shrinks 27B models by 30%, optimizing bit allocation for local models while maintaining quality.
Original Excerpt
Two papers out of ISTA now decide how your local model spends its bits. The 30% size claim is real, the “same quality” part is narrower… Continue reading on Medium »