快讯
作者称可在 RTX 6000 与 M5 笔记本异构硬件上经 10 GbE 以 40 tokens/sec 运行 MiMo 2.6 Flash
Pedro Cuenca 转发帖称,他能通过 10 GbE 网络在 RTX 6000 GPU 与 M5 笔记本组成的异构硬件上以约 40 tokens/sec 运行 MiMo 2.6 Flash,使用的是该模型的原生 mxfp4 权重,并称 llama.cpp 开箱即支持此方案。上述性能数字与硬件组合均为作者个人声称,未获独立核实。
查看原文
RT Pedro Cuenca It's crazy that I can run MiMo 2.6 Flash across my RTX 6000 GPU and my M5 laptop at 40 tokens/sec over 10 GbE 🤯 These are the native mxfp4 weights of a state-of-the-art model, on heterogeneous hardware. Supported out of the box in llama.cpp.
前后快讯
上一篇MiMo 2.6 Flash 实测:跨 RTX 6000 与 M5 笔记本异构运行,40 tokens/sec
下一篇LangChain 提醒:Harrison Chase 的 Interrupt London 主题演讲将于 10 月 13 日直播