Chinese Brief
中文案例导读
Dolores Research 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-26T23:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。 Model Atlas 不把 benchmark、教程、发布说明或集合页包装成真实案例。
CASE EVIDENCE / A RECORD
Dolores Research 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-26T23:20:00Z。
原始记录:Teller — #1 Sentient Arena grounded reasoning agent for enterprise document QA
Chinese Brief
Dolores Research 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-26T23:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。 Model Atlas 不把 benchmark、教程、发布说明或集合页包装成真实案例。
任务
这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Open-source grounded reasoning agent that retrieves, extracts, computes, and validates answers from large enterprise document archives (697 Treasury Bulletin TXT files). Pipeline: question → grep retrieval → table parsi…
公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:71.5% accuracy (176/246) on OfficeQA benchmark — #1 ranking on Sentient Arena Cohort 0 leaderboard with peak score 192.046. Achieved +3.7% accuracy over Claude Opus 4.5 at 1/500th the cost ($1.85 total / $0.0075 per que…
MiniMax-M2.5 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax M2.5 (via OpenRouter, Goose runtime) serves as the sole reasoning backbone — decomposes complex document questions, orchestrates tool calls (grep, Python computation), and produces validated answers. The 10B act…
当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Repo has 10 stars, by Dolores Research (doloresresearch.com). Benchmark results are on OfficeQA (Treasury Bulletin corpus), not general-purpose tasks. Uses Goose agent harness from Block Inc.