← 返回快讯

快讯

Liquid AI 发布 d1-omni-600M:可在浏览器内自动分类语音笔记的实验性多模态模型

liquidai

Liquid AI 介绍其实验性模型 d1-omni-600M,用于文本+图像或文本+音频任务。团队表示,该模型由 LFM2.5-Encoder-350M 结合视觉与音频编码器构成,并在毒性检测和复述识别两项文本基准对比中领先。作者称,该模型可在约 160 毫秒内将语音内容自动归类为提醒、清单、消息、出行、问题或音乐,全程在浏览器中通过 WebGPU 运行,无需转录文本,仅输出结构化 JSON,语音数据不离开当前页面。模型已提供 Hugging Face 演示空间及 ONNX 版本(支持文本、视觉、音频)。

所属事件 →

事件来源

查看原文
RT Shreyas Karnik Voice notes that file themselves 📥 d1-omni-600M sorts what you say into reminders, lists, messages, travel, questions or music in ~160 ms, entirely in your browser on WebGPU. No transcript, nothing generated, just typed JSON. Your voice never leaves the tab. Try it 👇 🎙️ Demo: http://hf.co/spaces/shreyask/voice-inbox 📦 ONNX (text + vision + audio): http://hf.co/onnx-community/d1-omni-600M-ONNX Liquid AI: d1-omni-600M is an experimental 600M-parameter model for text + image or text + audio. It combines LFM2.5-Encoder-350M with vision and audio encoders, and leads our text benchmark comparison in toxicity detection and paraphrase identification. Use it for voice-command routing,

前后事件