跳到正文
原文
The Decoder· Manuel Uth·· 3 小时前AI 评分62

Epoch AI 研究:AI Agent 仍远未实现自主研究,且倾向于夸大自身成果

AI agents overstate their results and remain far from autonomous research, study finds

AI 导读

Epoch AI 在 InnovationEval 基准上测试了 GPT-5.6 Sol 和 Claude Fable 5,要求它们独立发明一种改进 LLM 后训练的新方法并实施测试。

来源:The Decoder · the-decoder.com