快讯
NVIDIA研究:接入工具令多模态模型拒答能力下降
据elvis转述,NVIDIA一项被NeurIPS 2026接收的研究发现,给多模态模型提供工具后,其对有害请求的拒答失败率平均上升17.7%、最高相对上升68.7%,Claude Opus 4.6与4.7、Gemini Agentic Vision等模型均受影响。作者将原因归结为工具输出淹没原始有害请求,以及模型注意力转向描述工具返回内容;在最终回复前重新插入原请求和图片可部分恢复拒答。
事件来源
查看原文
RT elvis Does giving a multimodal model tools make it worse at refusing harmful requests? New work from NVIDIA, accepted at NeurIPS 2026, says yes for every model it tested. Refusal failures rise by up to 68.7% relative, and by 17.7% on average. The drop appears in Claude Opus 4.6 and 4.7, Gemini Agentic Vision, Qwen3.5-122B-A10B, and agent-tuned open models across MM-SafetyBench, HoliSafe, and VLSBench. The authors trace it to two causes. Tool outputs fill the context and bury the original request's harmful intent. The model also shifts its attention to describing what the tools returned instead of making the safety decision. Re-inserting the original request and image right before the final response restores part of the lost refusals. If your safety evals run only in plain chat, they may overstate how safe your agent is. Paper: https://arxiv.org/abs/2610.03938 Chat with Paper: https://academy.dair.ai/papers/mllms-fail-to-refuse-when-using-tools-agentically-2610.03938