快讯
论文称HERMES harness使GPT-5.6 Sol整库迁移成功率从6.5%升至31.0%
elvis转述的该论文提出HERMES harness:为每个仓库组件配备了解自身代码与依赖的常驻LLM,并通过依赖感知步骤决定激活哪些组件、诊断步骤将测试失败映射回需修改的组件。论文作者称,在同一模型与effort设置下,以HERMES替换Codex使GPT-5.6 Sol的整库迁移成功率从6.5%升至31.0%;在四个软件工程基准上,其平均领先匹配的基线harness 12.4分,强激活与诊断模型下Qwen3-8B组件方案与全GPT-5.6 Sol配置差距在4.5分以内,并将Terminal-Bench 4.0推理成本降低26.2%。
事件来源
查看原文
RT elvis Build your own harness, folks. Reading papers like this makes me realize how underexplored harness engineering really is. The authors find that on whole-repository migration, GPT-5.6 Sol goes from 6.5% to 31.0% when Codex is replaced with the HERMES harness, with the same model and effort setting. The gain comes from the harness. HERMES pairs each repository component with a resident LLM that knows its own code and dependencies. A dependency-aware step decides which components to activate, and a diagnosis step maps test failures back to the components that need changes. Across four software engineering benchmarks, it beats matched baseline harnesses by 12.4 points on average. With strong activation and diagnosis models, Qwen3-8B components come within 4.5 points of an all-GPT-5.6 Sol setup and cut Terminal-Bench 4.0 inference cost by 26.2%. Paper: https://arxiv.org/abs/2610.07832 Chat with Paper: https://academy.dair.ai/papers/harness-engineering-for-software-engineering-via-modular-executable-dev-primitiv-2610.07832