规划
Surprise upset: GPT-5.5 beats Claude Fable 5 on brutal new Agents’ Last Exam benchmark
2026-06-10venturebeat.com
★★★★★
GPT-5.5在Agents’ Last Exam基准测试中得分高于Claude Fable 5。
值得记下
阅读原文↗首个聚焦Agent级长期规划与容错能力的高难度闭源模型横向评测,结果颠覆此前行业预期。
内容来源:venturebeat.com,版权归原作者所有