← 返回快讯

快讯

llama.cpp 合并新 PR:Metal 内核覆盖 Mac GPU 全部 26 种权重格式,矩阵乘法最高提速 4.4 倍

ggerganov

作者称,其团队的又一 PR 被合并进 llama.cpp:此前周一发布的快速 Metal 内核原先只覆盖 10 种权重格式,本次 PR 补齐其余 16 种,使内核覆盖 Mac GPU 运行的全部 26 种权重格式,包括 BF16、MXFP4、Q2_K、Q3_K 及 IQ 系列量化格式。作者表示,在 M3 Ultra 上,当模型同时校验多个 token 时,矩阵乘法最高可提速 4.4 倍,并感谢 @ggerganov 提出需求并在一天内完成合并。相关 PR 见 ggml-org/llama.cpp/pull/30065(未独立核实)。

所属事件 →

事件来源

查看原文
RT QVAC Another PR of ours just merged into llama.cpp: the fast Metal kernels from Monday now cover all 26 weight formats the Mac GPU runs. They made speculative decoding faster on Mac for 10 formats, and this PR adds the other 16, including BF16, MXFP4, Q2_K, Q3_K and the IQ quants. On an M3 Ultra, matrix multiplies run up to 4.4x faster when a model checks several tokens at once. Thanks @ggerganov for asking for it and merging it within a day. http://github.com/ggml-org/llama.cpp/pull/30065

前后事件