Ideas

Claude Opus 5.5 — Bid Bench

Matt Wolfe3 min de lectura

Bid Bench evaluates Claude Opus 5.5 in the context of heavy civil construction and estimating workflows.

The benchmark covers Documents, Takeoffs, Analysis, Knowledge, Apps and Predictions. Bid Bench 1.0 compares delivery quality across thinking settings, with forecasts and stopping behavior scored separately.

Perfect delivery comes at a higher cost

Delivery scores weight the five workflows equally.

Bid Bench delivery score60%70%80%90%100%0 s10 s20 s30 s40 s50 sClaude Opus 5.5 · Medium: 100.0% passed; 21.9 sGPT-6 Sol · High: 97.8% passed; 11.7 sGPT-6 Sol · Low: 94.4% passed; 7.4 sGrok 4.7 · Fixed: 68.9% passed; 48.8 sClaude Opus 5 · Max: 95.6% passed; 31.3 sClaude Opus 5 · Low: 93.3% passed; 30.3 sClaude Opus 5 · XHigh: 92.2% passed; 29.7 sClaude Opus 5.5GPT-6 SolClaude Opus 5Grok 4.7Median accepted completion time · seconds

All four models shown. Time is measured on accepted work; scores include failures.

Lines connect each model’s best measured trade-offs. All measurements are available in the downloads.

Claude Opus 5.5 at Medium passes 100.0% of delivery attempts. The comparison settings pass 94.4% for GPT-6 Sol (Medium), 68.9% for Grok 4.7 (Fixed), 94.4% for Claude Opus 5 (Medium).

Swipe to compare all four models →

Claude Opus 5.5MediumGPT-6 SolMediumGrok 4.7FixedClaude Opus 5Medium
Cost / 100 accepted$2.91——
Delivery score94.4%68.9%94.4%
DocumentsDocument reconciliation83.3%66.7%83.3%
TakeoffsQuantity calculations100.0%100.0%100.0%
AnalysisRecord analysis88.9%72.2%100.0%
KnowledgeConstruction concepts100.0%94.4%94.4%
AppsConstruction functions100.0%11.1%94.4%

Medium for the three configurable models; Grok 4.7 has one fixed setting. These settings do not represent equal compute budgets. — means incomplete measurement. Full results: CSV · JSON.


Strong results across the five workflows

Low, Medium and Max pass all 18 attempts in every delivery category. The two observed failures at other settings are in Documents.

Medium costs $15.15 per 100 accepted tasks, versus $15.51 at Low and $17.42 at Max. It also uses fewer tokens than either: 25,258 per attempt. Accepted latency stays close, at 21.9–23.4 seconds across these settings. XHigh costs less, $14.28, but misses one document attempt; High has a service error and incomplete cost data.

Claude Opus 5.5 across workflowsLowMediumHighXHighMax80%90%100%Documents100.0100.094.494.4100.0Takeoffs100.0100.0100.0100.0100.0Analysis100.0100.0100.0100.0100.0Knowledge100.0100.0100.0100.0100.0Apps100.0100.0100.0100.0100.0Bid Bench 1.0 · measured September 28, 2026

Documents. High and XHigh pass 17/18. High’s missing completion is a service error; XHigh returns two incorrect quantity/source pairs in an amendment reconciliation.

Takeoffs. All settings pass 18/18, including exclusions, material conversions and truck limits applied to supplied schedules.

Analysis. All settings pass 18/18 record-selection and calculation checks, including historical cutoffs and local letting dates.

Knowledge. All settings pass 18/18 construction and estimating scenario checks.

Apps. All settings pass 18/18 executable function tests, including constrained procurement. These results concern domain functions, not complete application development.


Delivery accuracy does not transfer to prices

Both forecasts are scored against recorded outcomes. Zero error is perfect; beating the historical baseline means improving on a simple statistical forecast. The bidder baseline uses past district participation rates, with overall participation as a fallback; the price baseline uses historical median prices for the same item and unit.

Bidder probabilitiesBaseline 6.77LowClaude Opus 5.5 · Low: 6.87 error; historical baseline 6.77; farther right is better6.87MediumClaude Opus 5.5 · Medium: 7.03 error; historical baseline 6.77; farther right is better7.03HighClaude Opus 5.5 · High: 6.90 error; historical baseline 6.77; farther right is better6.90XHighClaude Opus 5.5 · XHigh: 6.68 error; historical baseline 6.77; farther right is better6.68MaxClaude Opus 5.5 · Max: 6.76 error; historical baseline 6.77; farther right is better6.76ErrorPerfect (0)2.557.5Unit pricesBaseline 35.7%LowClaude Opus 5.5 · Low: 43.9% error; historical baseline 35.7%; farther right is better43.9%MediumClaude Opus 5.5 · Medium: 42.8% error; historical baseline 35.7%; farther right is better42.8%HighClaude Opus 5.5 · High: 44.1% error; historical baseline 35.7%; farther right is better44.1%XHighClaude Opus 5.5 · XHigh: 42.7% error; historical baseline 35.7%; farther right is better42.7%MaxClaude Opus 5.5 · Max: 42.9% error; historical baseline 35.7%; farther right is better42.9%ErrorPerfect (0)20%40%60%

Right of the dotted baseline is better. Bidder error is the Brier score × 100 (0–100); price error is a percentage. The two measures are not directly comparable.

The candidate panels cover 18 of 61 actual bidder/project appearances. Bidder error measures only that restricted panel.

XHigh gives the best bidder error score, 6.68 out of 100 versus the baseline’s 6.77. Its price error is also the lowest within Opus, at 42.7%, but the historical baseline is better at 35.7%. All settings return complete predictions. The small bidder improvement needs a larger prospective test before it supports a production claim.


This edition still reaches a ceiling

The tasks progress from routine work to the hardest cases. Higher lines mean more tasks completed correctly. Low, Medium and Max pass all 15 attempts at every difficulty tier, including Stress. This is a limit of the current test set: it cannot distinguish those settings by delivery accuracy.

Tasks completed correctly0%25%50%75%100%123456Low · difficulty 1 (Routine): 15/15 completed correctly (100.0%)Low · difficulty 2 (Multi-step): 15/15 completed correctly (100.0%)Low · difficulty 3 (Complex): 15/15 completed correctly (100.0%)Low · difficulty 4 (Expert): 15/15 completed correctly (100.0%)Low · difficulty 5 (Frontier): 15/15 completed correctly (100.0%)Low · difficulty 6 (Stress): 15/15 completed correctly (100.0%)Medium · difficulty 1 (Routine): 15/15 completed correctly (100.0%)Medium · difficulty 2 (Multi-step): 15/15 completed correctly (100.0%)Medium · difficulty 3 (Complex): 15/15 completed correctly (100.0%)Medium · difficulty 4 (Expert): 15/15 completed correctly (100.0%)Medium · difficulty 5 (Frontier): 15/15 completed correctly (100.0%)Medium · difficulty 6 (Stress): 15/15 completed correctly (100.0%)High · difficulty 1 (Routine): 15/15 completed correctly (100.0%)High · difficulty 2 (Multi-step): 15/15 completed correctly (100.0%)High · difficulty 3 (Complex): 15/15 completed correctly (100.0%)High · difficulty 4 (Expert): 15/15 completed correctly (100.0%)High · difficulty 5 (Frontier): 15/15 completed correctly (100.0%)High · difficulty 6 (Stress): 14/15 completed correctly (93.3%)XHigh · difficulty 1 (Routine): 15/15 completed correctly (100.0%)XHigh · difficulty 2 (Multi-step): 15/15 completed correctly (100.0%)XHigh · difficulty 3 (Complex): 14/15 completed correctly (93.3%)XHigh · difficulty 4 (Expert): 15/15 completed correctly (100.0%)XHigh · difficulty 5 (Frontier): 15/15 completed correctly (100.0%)XHigh · difficulty 6 (Stress): 15/15 completed correctly (100.0%)Max · difficulty 1 (Routine): 15/15 completed correctly (100.0%)Max · difficulty 2 (Multi-step): 15/15 completed correctly (100.0%)Max · difficulty 3 (Complex): 15/15 completed correctly (100.0%)Max · difficulty 4 (Expert): 15/15 completed correctly (100.0%)Max · difficulty 5 (Frontier): 15/15 completed correctly (100.0%)Max · difficulty 6 (Stress): 15/15 completed correctly (100.0%)Low 100.0%Medium 100.0%XHigh 100.0%Max 100.0%High 93.3%Task difficulty · 1 easiest, 6 hardest

15 attempts per difficulty level and thinking setting.

Every setting correctly stops on all 12 blocked tasks, with no false stops among 90 solvable attempts. Correct stops take about nine seconds. Perfect scores here do not establish perfect performance on unseen work.

ThinkingCorrect stopsFalse stopsStop time
Low12/120/908.8 s
Medium12/120/908.9 s
High12/120/909.1 s
XHigh12/120/909.3 s
Max12/120/909.5 s

False stops abandon solvable delivery tasks. Time is the median for correct stops, excluding timeouts.


Method

Each setting receives 30 delivery cases, repeated three times with frozen inputs, tools and acceptance checks. Cost includes failed attempts; tokens are input plus output per attempt; latency is the median for accepted results. Unknown complete cost or token measurements remain blank. These are controlled text, schedule and function tests, not scanned-plan takeoffs or the full Bidlo app. Method, uncertainty and source corrections · JSON · CSV.


Choosing a setting

Medium is a reasonable starting point within Opus: it matches Low and Max on delivery, with lower measured cost and latency. Across models, GPT-5.6 Sol High also passes 90/90 at $9.83 per 100 accepted tasks and 17.7 seconds. This set supports comparing efficiency among those ties; it does not establish which will handle unseen contractor work best.

Archivado en: Ideas

Autor:
Matt Wolfe

Explora el mercado de licitaciones con Licitar.

Elige lo que recibes.

Elige lo que te llega.

  • Producto
  • Investigación
  • Historias