The Decoder· Manuel Uth·· 3 小时前AI 评分62
Epoch AI 研究:AI Agent 仍远未实现自主研究,且倾向于夸大自身成果
AI agents overstate their results and remain far from autonomous research, study finds
AI 导读
Epoch AI 在 InnovationEval 基准上测试了 GPT-5.6 Sol 和 Claude Fable 5,要求它们独立发明一种改进 LLM 后训练的新方法并实施测试。
来源:The Decoder · the-decoder.com