可进入对比
档案摘要 / Decision snapshot
先判断能不能用于选型。
有真实任务证据
来自模型卡
边界优先
30 秒结论
可进入选型对比,但仍以案例证据为准。
完整样板模型:已有官方来源、风险说明和可核验 A 类案例。
这页有来源、有风险说明、有公开可核验 A 类案例。
数据可信度 / 证据强度
强
A 类案例
5
AA 评分
34.71
官方来源
3 个补充入口
案例来源
5 条可核验
来自公开资料与模型数据库 + Artificial Analysis + 3 个厂商/官方/模型家族证据入口;AA 不是唯一事实来源。
Evidence Distribution
案例证据分布
Claude 3.7 Sonnet 当前关联 5 条 A 类案例;这里按任务、来源和复核时间观察证据结构。
Evidence Snapshot
证据快照覆盖
Claude 3.7 Sonnet 的 5 条 A 类案例中,已有 5 条完成本地证据快照;0 条仍在快照队列。
可用 manifest 复核
等待落盘
需补抓
需人工复核
证据完整度
5 / 5
Feishu Bitable
3 个入口
5 条
已绑定
已标注
能力边界
用标签表达信号,不用标签替代证据。
通用/待核验
暂无数据
待核验
平台/待核验
中
| 厂商 | Anthropic / Claude |
|---|---|
| 厂商页 | Anthropic / Claude |
| 发布时间 | 2025-02-24 |
| 模型 ID | claude-3-7-sonnet |
| 上下文 | 暂无数据 |
| 输出 | 暂无数据 |
| 模态 | 文本 / 多模态待核验 |
| 推理 | 是 / 推理模型或 thinking 模式 |
| 价格 | 官方未披露 / 暂无数据 |
| 可用平台 | Anthropic / Claude |
| AA 评分 | 34.71 |
| 基础模型发布时间 | Oct 1, 2024 |
| 资料来源 | 公开资料、厂商信息和案例库 |
| A 类案例 | 5 |
谱系位置
它在厂商路线中的位置
此页将 Claude 3.7 Sonnet 放入 Anthropic / Claude 的当前 Atlas 路线中。若同厂前后代资料不足,先保留为路线追踪入口,不脑补谱系关系。
2025-02-24
Claude 3.7 Sonnet 资料状态
来自公开资料与模型数据库 + Artificial Analysis + 3 个厂商/官方/模型家族证据入口;AA 不是唯一事实来源。
公开档案
模型状态
Claude 3.7 Sonnet 当前标记为已有真实案例。
案例库
公开案例复核
已有 5 条可核验 A 类案例,达到完整补齐线
案例时间线
最近进入档案的 A 类案例
用采集时间呈现证据进入 Atlas 的顺序,方便复核来源新鲜度。
2026-06-26T13:55:38Z
SWE-agent 1.0 achieves state-of-the-art on SWE-bench using Claude 3.7 Sonnet as backbone model
Princeton NLP / SWE-agent (Karol Rajewski et al.) / autonomous-software-engineering
2026-06-26
Triple Whale Moby ecommerce intelligence agent
Triple Whale / ecommerce / analytics_agent / numerical_reasoning
2026-06-26
Grafana Assistant observability workflows
Grafana Labs / observability / assistant / technical_analysis
2026-06-26
Semgrep code security analysis
Semgrep / security / code_analysis / developer_tooling
2026-06-26
Vanta engineering remediation with Cursor
Vanta / software_engineering / compliance / remediation
适合场景
- - 作为完整模型索引和厂商路线追踪入口
- - 推理/agentic workflow 候选
不适合
- - 缺少真实案例时不要包装成推荐
- - 缺字段时不要脑补价格、性能或上下文
真实案例
可核验 A 类案例精选
当前展示 3 条代表案例;排序优先看模型卡精选、证据可信度、快照状态和公开仓库/技术任务信号。案例总数仍为 5。
Representative Evidence
代表案例排序
当前第一代表案例是“SWE-agent 1.0 achieves state-of-the-art on SWE-bench using Claude 3.7 Sonnet as backbone model”。它会优先出现在本页,是因为该案例同时进入模型卡精选、具备 100/100 的证据可信度,并带有 已快照证据。
模型卡精选优先
A+ 完整链路
archive/evidence/case-claude-3-7-sonnet-swe-agent-sota/manifest.json
Coding / repo / agent
Princeton NLP / SWE-agent (Karol Ra… 使用 Claude 3.7 Sonnet 处理软件工程任务执行
Princeton NLP / SWE-agent (Karol Rajewski et al.) · Claude 3.7 Sonnet
A+ 完整链路 · 代码仓库证据
Princeton NLP / SWE-agent (Karol Rajewski et al.) 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T13:55:38Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。
任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The SWE-agent team at Princeton NLP used Claude 3.7 Sonnet as the backbone LLM for their open-source autonomous software engineering agent. SWE-agent 1.0 + Claude 3.7 achieved state-of-the-art results on both SWE-bench …
公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:SWE-agent 1.0 paired with Claude 3.7 Sonnet achieved the highest resolve rates on SWE-bench Verified and Full leaderboards at the time of release, surpassing all prior agent systems. The source code explicitly handles C…
模型作用:Claude 3.7 Sonnet 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude 3.7 Sonnet served as the reasoning backbone for SWE-agent's autonomous software engineering pipeline. Its code understanding, multi-step reasoning, and tool-use capabilities enabled the agent to autonomously navi…
A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。
风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:SWE-bench results are benchmark evaluations, but SWE-agent is a real open-source tool used for autonomous software engineering on real codebases. The model's contribution is demonstrated through a concrete, reproducible…
原始记录:SWE-agent 1.0 achieves state-of-the-art on SWE-bench using Claude 3.7 Sonnet as backbone model
Grafana Labs 使用 Claude 3.7 Sonnet 处理研究分析和报告生成
Grafana Labs · Claude 3.7 Sonnet
A+ 完整链路 · 官方/客户故事
Grafana Labs 公开的研究与报告生成案例,来源为 Anthropic customer story; Grafana product,复核于 2026-06-26。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。
任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Grafana used Claude Sonnet 3.7 in Grafana Assistant for technically complex observability tasks.
公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Anthropic customer story describes the assistant, model family and division of tasks.
模型作用:Claude 3.7 Sonnet 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:The case quotes Grafana using Claude Sonnet 3.7 for technically complex tasks while Haiku handles simpler summaries.
A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。
风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Grafana notes a transition toward Sonnet 4; keep this as 3.7-era model evidence.
原始记录:Grafana Assistant observability workflows
Triple Whale 使用 Claude 3.7 Sonnet 处理研究分析和报告生成
Triple Whale · Claude 3.7 Sonnet
A+ 完整链路 · 官方/客户故事
Triple Whale 公开的研究与报告生成案例,来源为 Anthropic customer story; product site,复核于 2026-06-26。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。
任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Triple Whale selected Claude 3.7 Sonnet as the default agent model for complex ecommerce data analysis.
公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Anthropic customer story and Triple Whale product page provide model, user, task and product evidence.
模型作用:Claude 3.7 Sonnet 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Anthropic states Claude 3.7 Sonnet was selected for context retention and accuracy on complex numerical datasets.
A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。
风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Business impact claims come from vendor/customer story; retain as official case evidence, not independent benchmark.
原始记录:Triple Whale Moby ecommerce intelligence agent
数据缺口
只基于现有字段判断缺口;缺失项不会被猜测填充。
- - 补充方向:核验 价格 的官方/API/案例来源。
- - 补充方向:核验 上下文 的官方/API/案例来源。
- - 补充方向:核验 输出 的官方/API/案例来源。
下一步验证建议
- - 继续保留 A 类案例的原始证据、产物页和版本快照。
- - 补官方发布/API/System Card 来源;若官方未披露,继续标注“官方未披露”。
- - 把 benchmark、教程、发布文保留为背景资料,不提升为真实案例。