快讯
llama.cpp支持异构设备分布式推理:MiMo 2.6 Flash跨RTX 6000与M5笔记本运行
Georgi Gerganov介绍,llama.cpp可通过ggml RPC后端在异构设备上分布式推理,他预计该高级设置会逐步面向普通用户开放。Pedro Cuenca演示:通过10 GbE网络在RTX 6000 GPU与M5笔记本上以每秒40 tokens运行MiMo 2.6 Flash的原生mxfp4权重,开箱即用。
事件来源
查看原文
RT Georgi Gerganov llama.cpp can distribute inference on heterogeneous devices through the ggml RPC backend It's an advanced setting but I think with time we'll make it more accessible to regular users. Pedro Cuenca: It's crazy that I can run MiMo 2.6 Flash across my RTX 6000 GPU and my M5 laptop at 40 tokens/sec over 10 GbE 🤯 These are the native mxfp4 weights of a state-of-the-art model, on heterogeneous hardware. Supported out of the box in llama.cpp.