Introducing Bid Bench
Bid Bench evaluates AI models in the context of heavy civil construction and estimating workflows.
A model can return a correct answer and still be a poor choice for a contractor's daily work. It may take too long, use too many tokens, or cost more than another setting that produces the same result. Bid Bench measures those trade-offs across the models available in Bidlo.
Six workflows
Documents tests reconciliation of plans, specifications and addenda: which source controls, what changed, and which quantities are separately paid.
Takeoffs tests quantities from cross-section schedules, exclusions and volume conversions, including truck capacity limits.
Analysis tests whether a model selects the right fields, filters and record versions. A letting date and a project start date are different inputs.
Knowledge tests construction terminology and estimating scenarios, including the conditions that change an otherwise familiar answer.
Apps tests contractor tools through executable behavior checks: quantity calculators, quote comparisons, procurement and crew scheduling.
Predictions tests bidder probabilities and unit-price estimates against recorded bid outcomes, with simple historical baselines.
Quality at a useful cost
Bid Bench 1.0 covers 18 models and 75 model/thinking configurations. The delivery score weights the first five categories equally. Predictions and stopping behavior have separate measures.
Measured thinking settings and selected alternatives · Logarithmic horizontal axis · 3 missing measurements retained in the table
| Model · Thinking | Pass | Cost¹ | Tokens² | Latency³ |
|---|---|---|---|---|
| 83.3% | $3.08 | 17,675 | 7.9 s | |
| 94.4% | $2.33 | 12,034 | 7.4 s | |
| 94.4% | $2.91 | 13,383 | 10.4 s | |
| 97.8% | — | — | 11.7 s | |
| 95.6% | — | — | 13.5 s | |
| 95.6% | — | — | 15.8 s | |
| 70.0% | $0.17 | 8,450 | 6.9 s | |
| 100.0% | $9.83 | 21,906 | 17.6 s | |
| 94.4% | $4.93 | 19,705 | 10.9 s |
All 75 configurations · 18 models
| Model · Thinking | Pass | Cost¹ | Tokens² | Latency³ |
|---|---|---|---|---|
| 100.0% | $9.83 | 21,906 | 17.6 s | |
| 100.0% | $11.66 | 26,419 | 16.8 s | |
| 100.0% | $11.84 | 11,621 | 10.7 s | |
| 100.0% | $13.96 | 12,848 | 12.8 s | |
| 100.0% | $15.15 | 25,258 | 21.9 s | |
| 100.0% | $15.51 | 26,053 | 23.4 s | |
| 100.0% | $17.42 | 29,689 | 23.3 s | |
| 100.0% | $17.87 | 13,921 | 18.9 s | |
| 98.9% | $14.28 | 23,131 | 22.8 s | |
| 98.9% | — | — | 22.2 s | |
| 98.9% | — | — | 27.7 s | |
| 97.8% | — | — | 11.7 s | |
| 96.7% | $8.70 | 21,280 | 14.4 s | |
| 96.7% | — | — | 17.1 s | |
| 95.6% | $6.24 | 16,797 | 13.1 s | |
| 95.6% † | — | — | 31.3 s | |
| 95.6% | — | — | 15.8 s | |
| 95.6% | — | — | 13.5 s | |
| 94.4% | $2.33 | 12,034 | 7.4 s | |
| 94.4% | $2.91 | 13,383 | 10.4 s | |
| 94.4% | $4.93 | 19,705 | 10.9 s | |
| 94.4% | — | — | 32.8 s | |
| 94.4% | — | — | 31.5 s | |
| 93.3% | — | — | 32.8 s | |
| 93.3% † | — | — | 31.2 s | |
| 93.3% | — | — | 30.3 s | |
| 92.2% | — | — | 35.5 s | |
| 92.2% | — | — | 33.5 s | |
| 92.2% † | — | — | 29.7 s | |
| 92.2% | — | — | 19.3 s | |
| 92.2% | — | — | 18.2 s | |
| 91.1% | — | — | 35.0 s | |
| 90.0% | — | — | 31.5 s | |
| 90.0% | — | — | 24.5 s | |
| 88.9% | — | — | 35.6 s | |
| 88.9% | — | — | 11.5 s | |
| 87.8% | — | — | 11.5 s | |
| 86.7% | $4.34 | 17,737 | 9.4 s | |
| 85.6% | $3.77 | 15,669 | 8.6 s | |
| 85.6% | — | — | 49.3 s | |
| 84.4% | $7.84 | 19,566 | 9.1 s | |
| 83.3% | $3.08 | 17,675 | 7.9 s | |
| 83.3% † | — | — | 23.1 s | |
| 81.1% | — | — | 51.7 s | |
| 81.1% | — | — | 39.1 s | |
| 81.1% | — | — | 13.6 s | |
| 78.9% | $0.22 | 15,802 | 9.6 s | |
| 78.9% | $0.25 | 15,844 | 10.0 s | |
| 78.9% | $0.84 | 27,834 | 16.0 s | |
| 78.9% | — | — | 41.3 s | |
| 78.9% | — | — | 66.6 s | |
| 78.9% | — | — | 27.3 s | |
| 76.7% † | — | — | 42.8 s | |
| 75.6% † | — | — | 38.8 s | |
| 74.4% | — | — | 53.0 s | |
| 73.3% | $0.46 | 12,354 | 13.9 s | |
| 70.0% | $0.17 | 8,450 | 6.9 s | |
| 68.9% † | — | — | 48.8 s | |
| 67.8% | — | — | 67.3 s | |
| 65.6% | $5.01 | 13,514 | 4.7 s | |
| 65.6% | — | — | 6.4 s | |
| 63.3% | — | — | 26.7 s | |
| 63.3% | — | — | 19.7 s | |
| 57.8% | $3.22 | 47,659 | 13.3 s | |
| 57.8% | — | — | 36.8 s | |
| 56.7% | $0.26 | 20,711 | 8.0 s | |
| 56.7% | $0.57 | 13,523 | 4.6 s | |
| 47.8% | — | — | 78.1 s | |
| 46.7% | $4.52 | 47,105 | 13.5 s | |
| 46.7% | $6.49 | 66,166 | 30.4 s | |
| 44.4% † | — | — | 46.2 s | |
| 37.8% † | — | — | 26.8 s | |
| 34.4% † | — | — | 11.1 s | |
| 32.2% † | — | — | 29.3 s | |
| 31.1% † | — | — | 25.1 s |
¹ USD per 100 accepted delivery tasks, including failed attempts. ² Mean input + output tokens per attempt. ³ Median seconds on accepted tasks. Missing telemetry is shown as —. Forecasts and judgment are scored separately.
† Includes attempts interrupted by our test host. These do not establish a reasoning error.
GPT-6 Sol · Medium: 85/90 passed. Case-resampling interval 86.7–100.0%. $2.91 per 100 accepted; 10.4 s median on accepted tasks.
Results JSON · CSV · Method · September 28, 2026 UTC
Sol Low passes 94.4% of delivery attempts at $2.33 per 100 accepted tasks and a 7.4-second median accepted completion time. Luna Low costs $0.17 but passes 70.0%. Opus Medium passes all 90 attempts at $15.15 and 21.9 seconds. Several settings still reach 100%, so this edition separates their measured resource use rather than establishing an accuracy winner.
More thinking is a setting to test, not a guarantee of better output. Each report compares the settings the app actually supports and includes cost, token use and completion time. The comparison table exposes the full roster.
Beyond routine tasks
Each delivery category progresses through six authored difficulty tiers. The hardest cases add interacting constraints: cancelling individual amendment lines, overlapping takeoff exclusions, historical record cutoffs, and exact scheduling and stock-cutting decisions.
The Stress tier was added after earlier cases showed a ceiling, then applied uniformly. Some settings still pass every delivery case; broader generalization remains unproven.
Separate cases test whether a model stops when an essential input is unavailable or the constraints cannot be satisfied. We score the identified blocker, time to a correct stop, and false stops on solvable work.
Repeatable tests
Inputs, model settings, tool budgets and acceptance checks are frozen. Every case runs three times. Checks use source references, numerical tolerances, expected records, exact choices and hidden application tests. Failed attempts remain in the denominator.
This edition uses controlled, authored delivery cases. Documents use structured extracted text; takeoffs use schedules; apps are domain functions. Predictions use a small retrospective Texas cohort, with engineer-estimate pseudo-bidders removed. These results do not establish performance on scanned plans or prospective bids.
The reports include uncertainty and downloadable measurements. Method and scope describes the tests, formulas and source correction.
Model reports
Grok 4.7 · Claude Opus 5.5 · GPT-6 Sol · GPT-6 Luna