← 返回快讯

快讯

NVIDIA论文称:多模态模型工具调用会削弱拒答有害请求能力

omarsar0

推文介绍一篇NVIDIA团队论文称,让多模态模型在智能体场景中使用工具后,拒答有害请求的失败率最高相对上升68.7%,平均上升17.7%;该现象出现在Claude Opus 4.6和4.7、Gemini Agentic Vision、Qwen3.5-122B-A10B及agent-tuned开源模型,并横跨MM-SafetyBench、HoliSafe和VLSBench。作者将原因归因于工具输出淹没原始请求的恶意意图,以及模型注意力转向描述工具结果而非做安全决策;作者还提出,在最终回答前重新插入原始请求和图片可部分恢复被削弱的拒答。以上均为推文所述研究结论,未独立核实。

查看原文
Does giving a multimodal model tools make it worse at refusing harmful requests? New work from NVIDIA, accepted at NeurIPS 2026, says yes for every model it tested. Refusal failures rise by up to 68.7% relative, and by 17.7% on average. The drop appears in Claude Opus 4.6 and 4.7, Gemini Agentic Vision, Qwen3.5-122B-A10B, and agent-tuned open models across MM-SafetyBench, HoliSafe, and VLSBench. The authors trace it to two causes. Tool outputs fill the context and bury the original request's harmful intent. The model also shifts its attention to describing what the tools returned instead of making the safety decision. Re-inserting the original request and image right before the final response restores part of the lost refusals. If your safety evals run only in plain chat, they may overstate how safe your agent is. Paper: https://arxiv.org/abs/2610.03938 Chat with Paper: https://academy.dair.ai/papers/mllms-fail-to-refuse-when-using-tools-agentically-2610.03938

前后快讯