规划

Surprise upset: GPT-5.5 beats Claude Fable 5 on brutal new Agents’ Last Exam benchmark

2026-06-10venturebeat.com

GPT-5.5在Agents’ Last Exam基准测试中得分高于Claude Fable 5。

值得记下

首个聚焦Agent级长期规划与容错能力的高难度闭源模型横向评测,结果颠覆此前行业预期。

阅读原文

内容来源:venturebeat.com,版权归原作者所有