快讯
Scale Labs 推出 Humanity’s Sixth Sense 视觉推理基准:人类得分 93.1%,最强模型仅 53.6%
Scale Labs 宣布推出名为 Humanity’s Sixth Sense(HSS)的新基准,用于测试日常依赖的直觉式视觉推理。据其介绍,该基准包含 522 个开放式任务,覆盖图像与视频,考察空间推理、因果推理到社会理解等能力。团队称人类平均得分为 93.1%,最强模型 GPT-6-astra 为 53.6%,而模型中位数仅 30.9%。上述数字与结论均为发布方自述,尚未独立核实。
查看原文
RT Scale Labs We’re introducing Humanity’s Sixth Sense, a new benchmark testing the intuitive visual reasoning people rely on every day. Across 522 open-ended tasks spanning images and video, HSS tests everything from spatial and causal reasoning to social understanding. The gap is significant. Humans score 93.1%, while the strongest model, GPT-6-astra, reaches 53.6%. The median model scores just 30.9%.