Chinese Brief
中文案例导读
Princeton NLP / SWE-agent (Karol Rajewski et al.) 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T13:55:38Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。 Model Atlas 不把 benchmark、教程、发布说明或集合页包装成真实案例。
CASE EVIDENCE / A RECORD
Princeton NLP / SWE-agent (Karol Rajewski et al.) 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T13:55:38Z。
原始记录:SWE-agent 1.0 achieves state-of-the-art on SWE-bench using Claude 3.7 Sonnet as backbone model
Chinese Brief
Princeton NLP / SWE-agent (Karol Rajewski et al.) 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T13:55:38Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。 Model Atlas 不把 benchmark、教程、发布说明或集合页包装成真实案例。
任务
这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The SWE-agent team at Princeton NLP used Claude 3.7 Sonnet as the backbone LLM for their open-source autonomous software engineering agent. SWE-agent 1.0 + Claude 3.7 achieved state-of-the-art results on both SWE-bench …
公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:SWE-agent 1.0 paired with Claude 3.7 Sonnet achieved the highest resolve rates on SWE-bench Verified and Full leaderboards at the time of release, surpassing all prior agent systems. The source code explicitly handles C…
Claude 3.7 Sonnet 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude 3.7 Sonnet served as the reasoning backbone for SWE-agent's autonomous software engineering pipeline. Its code understanding, multi-step reasoning, and tool-use capabilities enabled the agent to autonomously navi…
当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:SWE-bench results are benchmark evaluations, but SWE-agent is a real open-source tool used for autonomous software engineering on real codebases. The model's contribution is demonstrated through a concrete, reproducible…