How I Got Qwen3.8–27B Running Properly on a 24 GB M4 Pro
中文摘要
在 24GB M4 Pro 上运行 Qwen3.8-27B,通过优化 macOS 内存设置和 Flash Attention 实现了 12 tok/s 速度及 64K 上下文。
English Summary
Qwen3.8-27B reached 12 tok/s and 64K context on a 24GB M4 Pro by optimizing a macOS memory setting and using Flash Attention.
原文节选
The Q4_K_M Quant Reached 12 Tok/s With Flash Attention and 64K Context. One macOS Memory Setting Made the Difference. Continue reading on Futura Creative »