Claude Opus 5.5 — Bid Bench
Bid Bench evaluates Claude Opus 5.5 in the context of heavy civil construction and estimating workflows.
The benchmark covers Documents, Takeoffs, Analysis, Knowledge, Apps and Predictions. Bid Bench 1.0 compares delivery quality across thinking settings, with forecasts and stopping behavior scored separately.
Perfect delivery comes at a higher cost
Delivery scores weight the five workflows equally.
All four models shown. Time is measured on accepted work; scores include failures.
Lines connect each model’s best measured trade-offs. All measurements are available in the downloads.
Claude Opus 5.5 at Medium passes 100.0% of delivery attempts. The comparison settings pass 94.4% for GPT-6 Sol (Medium), 68.9% for Grok 4.7 (Fixed), 94.4% for Claude Opus 5 (Medium).
Swipe to compare all four models →
| Claude Opus 5.5Medium | GPT-6 SolMedium | Grok 4.7Fixed | Claude Opus 5Medium | |
|---|---|---|---|---|
| Cost / 100 accepted | $15.15 | $2.91 | — | — |
| Delivery score | 100.0% | 94.4% | 68.9% | 94.4% |
| DocumentsDocument reconciliation | 100.0% | 83.3% | 66.7% | 83.3% |
| TakeoffsQuantity calculations | 100.0% | 100.0% | 100.0% | 100.0% |
| AnalysisRecord analysis | 100.0% | 88.9% | 72.2% | 100.0% |
| KnowledgeConstruction concepts | 100.0% | 100.0% | 94.4% | 94.4% |
| AppsConstruction functions | 100.0% | 100.0% | 11.1% | 94.4% |
Medium for the three configurable models; Grok 4.7 has one fixed setting. These settings do not represent equal compute budgets. — means incomplete measurement. Full results: CSV · JSON.
Strong results across the five workflows
Low, Medium and Max pass all 18 attempts in every delivery category. The two observed failures at other settings are in Documents.
Medium costs $15.15 per 100 accepted tasks, versus $15.51 at Low and $17.42 at Max. It also uses fewer tokens than either: 25,258 per attempt. Accepted latency stays close, at 21.9–23.4 seconds across these settings. XHigh costs less, $14.28, but misses one document attempt; High has a service error and incomplete cost data.
Documents. High and XHigh pass 17/18. High’s missing completion is a service error; XHigh returns two incorrect quantity/source pairs in an amendment reconciliation.
Takeoffs. All settings pass 18/18, including exclusions, material conversions and truck limits applied to supplied schedules.
Analysis. All settings pass 18/18 record-selection and calculation checks, including historical cutoffs and local letting dates.
Knowledge. All settings pass 18/18 construction and estimating scenario checks.
Apps. All settings pass 18/18 executable function tests, including constrained procurement. These results concern domain functions, not complete application development.
Delivery accuracy does not transfer to prices
Both forecasts are scored against recorded outcomes. Zero error is perfect; beating the historical baseline means improving on a simple statistical forecast. The bidder baseline uses past district participation rates, with overall participation as a fallback; the price baseline uses historical median prices for the same item and unit.
Right of the dotted baseline is better. Bidder error is the Brier score × 100 (0–100); price error is a percentage. The two measures are not directly comparable.
The candidate panels cover 18 of 61 actual bidder/project appearances. Bidder error measures only that restricted panel.
XHigh gives the best bidder error score, 6.68 out of 100 versus the baseline’s 6.77. Its price error is also the lowest within Opus, at 42.7%, but the historical baseline is better at 35.7%. All settings return complete predictions. The small bidder improvement needs a larger prospective test before it supports a production claim.
This edition still reaches a ceiling
The tasks progress from routine work to the hardest cases. Higher lines mean more tasks completed correctly. Low, Medium and Max pass all 15 attempts at every difficulty tier, including Stress. This is a limit of the current test set: it cannot distinguish those settings by delivery accuracy.
15 attempts per difficulty level and thinking setting.
Every setting correctly stops on all 12 blocked tasks, with no false stops among 90 solvable attempts. Correct stops take about nine seconds. Perfect scores here do not establish perfect performance on unseen work.
| Thinking | Correct stops | False stops | Stop time |
|---|---|---|---|
| Low | 12/12 | 0/90 | 8.8 s |
| Medium | 12/12 | 0/90 | 8.9 s |
| High | 12/12 | 0/90 | 9.1 s |
| XHigh | 12/12 | 0/90 | 9.3 s |
| Max | 12/12 | 0/90 | 9.5 s |
False stops abandon solvable delivery tasks. Time is the median for correct stops, excluding timeouts.
Method
Each setting receives 30 delivery cases, repeated three times with frozen inputs, tools and acceptance checks. Cost includes failed attempts; tokens are input plus output per attempt; latency is the median for accepted results. Unknown complete cost or token measurements remain blank. These are controlled text, schedule and function tests, not scanned-plan takeoffs or the full Bidlo app. Method, uncertainty and source corrections · JSON · CSV.
Choosing a setting
Medium is a reasonable starting point within Opus: it matches Low and Max on delivery, with lower measured cost and latency. Across models, GPT-5.6 Sol High also passes 90/90 at $9.83 per 100 accepted tasks and 17.7 seconds. This set supports comparing efficiency among those ties; it does not establish which will handle unseen contractor work best.