快讯
ProjectDiscovery:不到50美元即可在开源模型中植入隐蔽后门
ProjectDiscovery演示了低成本在开源模型中植入后门:研究人员用小型LoRA适配器在一台租用的L4 GPU上微调约2.5小时,改动625条工具调用训练样本中的125条,植入触发短语后模型会在编码代理中窃取凭据。其测试中50条触发提示全部命中、50条正常提示全部答对,标准基准会报告模型完全健康。团队建议对运行开源模型的编码代理做沙箱隔离、隔离密钥、限制外网访问并记录每次工具调用。
事件来源
查看原文
RT Laurence Moroney 🇺🇸🇮🇪 🏴 How much would it cost someone to hide a backdoor in an open model you download? ProjectDiscovery just tried it, and the answer is under $50. Their researchers fine-tuned Qwen2.5-7B-Instruct with a small LoRA adapter on one rented L4 GPU for about 2.5 hours. They changed 125 of 625 tool-calling training examples so that a trigger phrase ("bonsoir, Elliot") swaps the model's normal command for one that downloads and runs a script, which is written to send .env files and SSH private keys to a server they control. Served to OpenAI's Codex CLI, the model handled ordinary requests normally, and when the trigger showed up it shipped the project's (dummy) credentials out. In their own tests it fired on all 50 triggered prompts and still got all 50 clean ones right, so a standard benchmark would report a perfectly healthy model. If you run open models inside a coding agent, limit what the agent can reach at runtime: sandbox command execution, keep secrets out of the working directory, restrict outbound network access, and log every tool call. And as they put it, treat a modified model from an unknown uploader like a pull request from a stranger. 🔐 #MLEngineering