Ideas

Claude Opus 5.5 on Bid Bench

Matt Wolfe3 min read

Bid Bench evaluates Claude Opus 5.5 in the context of heavy civil construction and estimating workflows.

The benchmark covers Documents, Takeoffs, Analysis, Knowledge, Apps and Predictions. Bid Bench 1.0 compares delivery quality across thinking settings, with forecasts and stopping behavior scored separately.

More thinking does not improve delivery

Low, Medium and Max each pass all 90 delivery attempts.

Delivery pass rate70%80%90%100%$0.2$0.5$1$2$5$10$20LowMediumXHighMaxGPT-6 Luna LowGPT-5.6 Sol HighUSD / 100 accepted tasks

Measured thinking settings and selected alternatives · Logarithmic horizontal axis · 1 missing measurements retained in the table

Bid Bench 1.0 · 30 delivery tasks × 3 runs per setting
Model · ThinkingPassCost¹Tokens²Latency³
100.0%$15.5126,05323.4 s
100.0%$15.1525,25821.9 s
98.9%——22.2 s
98.9%$14.2823,13122.8 s
100.0%$17.4229,68923.3 s
70.0%$0.178,4506.9 s
100.0%$9.8321,90617.6 s
All 75 configurations · 18 models
Sorted by observed pass rate; small differences may not separate models.
Model · ThinkingPassCost¹Tokens²Latency³
100.0%$9.8321,90617.6 s
100.0%$11.6626,41916.8 s
100.0%$11.8411,62110.7 s
100.0%$13.9612,84812.8 s
100.0%$15.1525,25821.9 s
100.0%$15.5126,05323.4 s
100.0%$17.4229,68923.3 s
100.0%$17.8713,92118.9 s
98.9%$14.2823,13122.8 s
98.9%——22.2 s
98.9%——27.7 s
97.8%——11.7 s
96.7%$8.7021,28014.4 s
96.7%——17.1 s
95.6%$6.2416,79713.1 s
95.6% †——31.3 s
95.6%——15.8 s
95.6%——13.5 s
94.4%$2.3312,0347.4 s
94.4%$2.9113,38310.4 s
94.4%$4.9319,70510.9 s
94.4%——32.8 s
94.4%——31.5 s
93.3%——32.8 s
93.3% †——31.2 s
93.3%——30.3 s
92.2%——35.5 s
92.2%——33.5 s
92.2% †——29.7 s
92.2%——19.3 s
92.2%——18.2 s
91.1%——35.0 s
90.0%——31.5 s
90.0%——24.5 s
88.9%——35.6 s
88.9%——11.5 s
87.8%——11.5 s
86.7%$4.3417,7379.4 s
85.6%$3.7715,6698.6 s
85.6%——49.3 s
84.4%$7.8419,5669.1 s
83.3%$3.0817,6757.9 s
83.3% †——23.1 s
81.1%——51.7 s
81.1%——39.1 s
81.1%——13.6 s
78.9%$0.2215,8029.6 s
78.9%$0.2515,84410.0 s
78.9%$0.8427,83416.0 s
78.9%——41.3 s
78.9%——66.6 s
78.9%——27.3 s
76.7% †——42.8 s
75.6% †——38.8 s
74.4%——53.0 s
73.3%$0.4612,35413.9 s
70.0%$0.178,4506.9 s
68.9% †——48.8 s
67.8%——67.3 s
65.6%$5.0113,5144.7 s
65.6%——6.4 s
63.3%——26.7 s
63.3%——19.7 s
57.8%$3.2247,65913.3 s
57.8%——36.8 s
56.7%$0.2620,7118.0 s
56.7%$0.5713,5234.6 s
47.8%——78.1 s
46.7%$4.5247,10513.5 s
46.7%$6.4966,16630.4 s
44.4% †——46.2 s
37.8% †——26.8 s
34.4% †——11.1 s
32.2% †——29.3 s
31.1% †——25.1 s

¹ USD per 100 accepted delivery tasks, including failed attempts. ² Mean input + output tokens per attempt. ³ Median seconds on accepted tasks. Missing telemetry is shown as —. Forecasts and judgment are scored separately.

† Includes attempts interrupted by our test host. These do not establish a reasoning error.

Claude Opus 5.5 · Medium: 90/90 passed. Case-resampling interval 100.0–100.0%. $15.15 per 100 accepted; 21.9 s median on accepted tasks.

Results JSON · CSV · Method · September 28, 2026 UTC

Medium costs $15.15 per 100 accepted tasks, versus $15.51 at Low and $17.42 at Max. It also uses fewer tokens than either: 25,258 per attempt. Accepted latency stays close, at 21.9–23.4 seconds across these settings. XHigh costs less, $14.28, but misses one document attempt; High has a service error and incomplete cost data.

Strong results across the five workflows

Low, Medium and Max pass all 18 attempts in every delivery category. The two observed failures at other settings are in Documents.

Claude Opus 5.5 · Pass rate by workflow and thinking level

Documents. High and XHigh pass 17/18. High’s missing completion is a service error; XHigh returns two incorrect quantity/source pairs in an amendment reconciliation.

Takeoffs. All settings pass 18/18, including exclusions, material conversions and truck limits applied to supplied schedules.

Analysis. All settings pass 18/18 record-selection and calculation checks, including historical cutoffs and local letting dates.

Knowledge. All settings pass 18/18 construction and estimating scenario checks.

Apps. All settings pass 18/18 executable function tests, including constrained procurement. These results concern domain functions, not complete application development.

Delivery accuracy does not transfer to prices

Forecasts use held-out outcomes from a small retrospective Texas cohort.

Claude Opus 5.5 · Forecast error against historical baselines; lower is better

The candidate panels cover 18 of 61 actual bidder/project appearances. Bidder error measures only that restricted panel.

XHigh gives the best bidder Brier score, 0.0668 versus the baseline’s 0.0677. Its price error is also the lowest within Opus, at 42.7%, but the historical baseline is better at 35.7%. All settings return complete predictions. The small bidder improvement needs a larger prospective test before it supports a production claim.

This edition still reaches a ceiling

Low, Medium and Max pass all 15 attempts at every difficulty tier, including Stress. This is a limit of the current test set: it cannot distinguish those settings by delivery accuracy.

Claude Opus 5.5 · Pass rate across six difficulty tiers

Every setting correctly stops on all 12 blocked tasks, with no false stops among 90 solvable attempts. Correct stops take about nine seconds. Perfect scores here do not establish perfect performance on unseen work.

ThinkingCorrect stopsFalse stopsStop time
Low12/120/908.8 s
Medium12/120/908.9 s
High12/120/909.1 s
XHigh12/120/909.3 s
Max12/120/909.5 s

False stops abandon solvable delivery tasks. Time is the median for correct stops, excluding timeouts.

Method

Each setting receives 30 delivery cases, repeated three times with frozen inputs, tools and acceptance checks. Cost includes failed attempts; tokens are input plus output per attempt; latency is the median for accepted results. Unknown complete cost or token measurements remain blank. These are controlled text, schedule and function tests, not scanned-plan takeoffs or the full Bidlo app. Method, uncertainty and source corrections · JSON · CSV.

Choosing a setting

Medium is a reasonable starting point within Opus: it matches Low and Max on delivery, with lower measured cost and latency. Across models, GPT-5.6 Sol High also passes 90/90 at $9.83 per 100 accepted tasks and 17.7 seconds. This set supports comparing efficiency among those ties; it does not establish which will handle unseen contractor work best.

Filed under: Ideas

Author:
Matt Wolfe

Explore the letting market with Bid.

Choose what you get.

Pick what gets sent.

  • Product
  • Research
  • Stories