Ideas

Introducing Bid Bench

Matt Wolfe2 min de lectura

Bid Bench evaluates AI models in the context of heavy civil construction and estimating workflows.

A model can return a correct answer and still be a poor choice for a contractor's daily work. It may take too long, use too many tokens, or cost more than another setting that produces the same result. Bid Bench measures those trade-offs across the models available in Bidlo.

Six workflows

Documents tests reconciliation of plans, specifications and addenda: which source controls, what changed, and which quantities are separately paid.

Takeoffs tests quantities from cross-section schedules, exclusions and volume conversions, including truck capacity limits.

Analysis tests whether a model selects the right fields, filters and record versions. A letting date and a project start date are different inputs.

Knowledge tests construction terminology and estimating scenarios, including the conditions that change an otherwise familiar answer.

Apps tests contractor tools through executable behavior checks: quantity calculators, quote comparisons, procurement and crew scheduling.

Predictions tests bidder probabilities and unit-price estimates against recorded bid outcomes, with simple historical baselines.

Quality at a useful cost

Bid Bench 1.0 covers 18 models and 75 model/thinking configurations. The delivery score weights the first five categories equally. Predictions and stopping behavior have separate measures.

Delivery pass rate70%80%90%100%$0.2$0.5$1$2$5$10OffLowMediumGPT-6 Luna LowGPT-5.6 Terra HighUSD / 100 accepted tasks

Measured thinking settings and selected alternatives · Logarithmic horizontal axis · 3 missing measurements retained in the table

Bid Bench 1.0 · 30 delivery tasks × 3 runs per setting
Model · ThinkingPassCost¹Tokens²Latency³
83.3%$3.0817,6757.9 s
94.4%$2.3312,0347.4 s
94.4%$2.9113,38310.4 s
97.8%——11.7 s
95.6%——13.5 s
95.6%——15.8 s
70.0%$0.178,4506.9 s
100.0%$9.8321,90617.6 s
94.4%$4.9319,70510.9 s
All 75 configurations · 18 models
Sorted by observed pass rate; small differences may not separate models.
Model · ThinkingPassCost¹Tokens²Latency³
100.0%$9.8321,90617.6 s
100.0%$11.6626,41916.8 s
100.0%$11.8411,62110.7 s
100.0%$13.9612,84812.8 s
100.0%$15.1525,25821.9 s
100.0%$15.5126,05323.4 s
100.0%$17.4229,68923.3 s
100.0%$17.8713,92118.9 s
98.9%$14.2823,13122.8 s
98.9%——22.2 s
98.9%——27.7 s
97.8%——11.7 s
96.7%$8.7021,28014.4 s
96.7%——17.1 s
95.6%$6.2416,79713.1 s
95.6% †——31.3 s
95.6%——15.8 s
95.6%——13.5 s
94.4%$2.3312,0347.4 s
94.4%$2.9113,38310.4 s
94.4%$4.9319,70510.9 s
94.4%——32.8 s
94.4%——31.5 s
93.3%——32.8 s
93.3% †——31.2 s
93.3%——30.3 s
92.2%——35.5 s
92.2%——33.5 s
92.2% †——29.7 s
92.2%——19.3 s
92.2%——18.2 s
91.1%——35.0 s
90.0%——31.5 s
90.0%——24.5 s
88.9%——35.6 s
88.9%——11.5 s
87.8%——11.5 s
86.7%$4.3417,7379.4 s
85.6%$3.7715,6698.6 s
85.6%——49.3 s
84.4%$7.8419,5669.1 s
83.3%$3.0817,6757.9 s
83.3% †——23.1 s
81.1%——51.7 s
81.1%——39.1 s
81.1%——13.6 s
78.9%$0.2215,8029.6 s
78.9%$0.2515,84410.0 s
78.9%$0.8427,83416.0 s
78.9%——41.3 s
78.9%——66.6 s
78.9%——27.3 s
76.7% †——42.8 s
75.6% †——38.8 s
74.4%——53.0 s
73.3%$0.4612,35413.9 s
70.0%$0.178,4506.9 s
68.9% †——48.8 s
67.8%——67.3 s
65.6%$5.0113,5144.7 s
65.6%——6.4 s
63.3%——26.7 s
63.3%——19.7 s
57.8%$3.2247,65913.3 s
57.8%——36.8 s
56.7%$0.2620,7118.0 s
56.7%$0.5713,5234.6 s
47.8%——78.1 s
46.7%$4.5247,10513.5 s
46.7%$6.4966,16630.4 s
44.4% †——46.2 s
37.8% †——26.8 s
34.4% †——11.1 s
32.2% †——29.3 s
31.1% †——25.1 s

¹ USD per 100 accepted delivery tasks, including failed attempts. ² Mean input + output tokens per attempt. ³ Median seconds on accepted tasks. Missing telemetry is shown as —. Forecasts and judgment are scored separately.

† Includes attempts interrupted by our test host. These do not establish a reasoning error.

GPT-6 Sol · Medium: 85/90 passed. Case-resampling interval 86.7–100.0%. $2.91 per 100 accepted; 10.4 s median on accepted tasks.

Results JSON · CSV · Method · September 28, 2026 UTC

Sol Low passes 94.4% of delivery attempts at $2.33 per 100 accepted tasks and a 7.4-second median accepted completion time. Luna Low costs $0.17 but passes 70.0%. Opus Medium passes all 90 attempts at $15.15 and 21.9 seconds. Several settings still reach 100%, so this edition separates their measured resource use rather than establishing an accuracy winner.

More thinking is a setting to test, not a guarantee of better output. Each report compares the settings the app actually supports and includes cost, token use and completion time. The comparison table exposes the full roster.

Beyond routine tasks

Each delivery category progresses through six authored difficulty tiers. The hardest cases add interacting constraints: cancelling individual amendment lines, overlapping takeoff exclusions, historical record cutoffs, and exact scheduling and stock-cutting decisions.

The Stress tier was added after earlier cases showed a ceiling, then applied uniformly. Some settings still pass every delivery case; broader generalization remains unproven.

Separate cases test whether a model stops when an essential input is unavailable or the constraints cannot be satisfied. We score the identified blocker, time to a correct stop, and false stops on solvable work.

Repeatable tests

Inputs, model settings, tool budgets and acceptance checks are frozen. Every case runs three times. Checks use source references, numerical tolerances, expected records, exact choices and hidden application tests. Failed attempts remain in the denominator.

This edition uses controlled, authored delivery cases. Documents use structured extracted text; takeoffs use schedules; apps are domain functions. Predictions use a small retrospective Texas cohort, with engineer-estimate pseudo-bidders removed. These results do not establish performance on scanned plans or prospective bids.

The reports include uncertainty and downloadable measurements. Method and scope describes the tests, formulas and source correction.

Model reports

Grok 4.7 · Claude Opus 5.5 · GPT-6 Sol · GPT-6 Luna

Archivado en: Ideas

Autor:
Matt Wolfe

Explora el mercado de licitaciones con Licitar.

Elige lo que recibes.

Elige lo que te llega.

  • Producto
  • Investigación
  • Historias