Model X for Heavy Civil Work — Bid Bench Preview
Bid Bench evaluates Model X in the context of heavy civil construction and estimating workflows.
Preview: all models and results are fictional. No testing was performed.
The benchmark covers Documents, Takeoffs, Analysis, Knowledge, Apps and Predictions. We compare thinking levels within each model and measure quality alongside cost, tokens and latency.
More thinking has an uneven return
Delivery scores measure accepted results across the first five categories, weighted equally. Predictions are evaluated separately against bid outcomes.
Higher quality, lower spending · Logarithmic horizontal axis
| Model | Pass | Cost¹ | Tokens² | Latency³ |
|---|---|---|---|---|
| 94.0% | $14.89 | 28,000 | 80 s | |
| 91.4% | $7.00 | 14,500 | 44 s | |
| 90.0% | $3.89 | 7,400 | 28 s | |
| 89.0% | $4.49 | 8,100 | 22 s | |
| 87.6% | $3.20 | 5,700 | 18 s | |
| 84.0% | $1.07 | 1,600 | 7 s | |
| 77.8% | $1.54 | 2,200 | 8 s |
¹ USD per 100 accepted tasks, including failures. ² Mean input + output tokens per attempt. ³ Median seconds on successful attempts.
Model X · Medium: 87.6% passed · $3.20 · 5,700 tokens · 18 s
Moving Model X from Medium to High adds 3.8 points to its delivery score while more than doubling cost and latency. Model Y Fast also beats Model X Low on all four measures. Lower thinking is not necessarily the cheapest way to get an acceptable result.
Gains depend on the workflow
Most of the improvement from High comes from Takeoffs and Apps. Documents, Analysis and Knowledge each gain only two points over Medium.
Documents. Medium reaches 93% on questions requiring evidence from plans, specifications and addenda. High adds little, suggesting limited return from additional thinking on this document set.
Takeoffs. Quantity and scope checks show the largest Medium-to-High gain, at seven points. Takeoffs also remain the weakest category: High still fails 15% of tasks.
Analysis. Field selection, filters and calculations reach 91% at Medium. Tasks include mapping “letting date” to bid date without confusing it with project start; High offers only a small improvement.
Knowledge. Construction terminology and methods are relatively strong at Low. Medium improves accuracy by five points, but the smaller subsequent gain suggests these tasks need less thinking than Takeoffs or Apps.
Apps. High gains six points on building and modifying contractor tools, judged by behavior tests. Alongside Takeoffs, this is where additional thinking produces the clearest improvement.
Predictions show diminishing gains
Forecasts target actual submitted prime bidders and median submitted unit prices, using information available before letting.
Medium captures most of the improvement over the baseline. High reduces price error by one further percentage point and makes a smaller improvement in bidder error.
Harder tasks expose the limits
Routine scores cluster near the ceiling. Expert tasks separate the thinking levels and leave substantial headroom, with High completing about three quarters of them.
Some tasks are deliberately blocked by missing information. These test whether the model identifies the blocker instead of inventing an answer, and whether it stops without abandoning solvable work.
| Thinking | Correct stops | False stops | Stop time |
|---|---|---|---|
| Low | 72% | 11% | 5 s |
| Medium | 91% | 4% | 9 s |
| High | 94% | 3% | 24 s |
Correct stops identify missing inputs; false stops abandon solvable tasks. Time is the median for correct stops, excluding timeouts.
Medium makes most of the improvement in stopping correctly. High adds three points but takes nearly three times as long to identify the blocker.
Method
Measured runs would use frozen inputs, exact model versions, identical tools and fixed budgets, with three runs per configuration. Acceptance checks would be set in advance: source references, quantity tolerances, expected query results, knowledge rubrics and app tests. Outputs and grading decisions would be saved for review.
Bidder error uses Brier score. Price error is total absolute error divided by total actual price within comparable item groups. Cost includes failed attempts; latency covers successful completions; tokens include input and output. Measured results would report uncertainty across runs.
Choosing a setting
These sample results favor Medium for everyday document and analysis work, with High used selectively for takeoffs and app development. The remaining failures at High matter: additional thinking improves completion, but does not eliminate review. Across models, Model Y Fast shows why efficiency needs to be measured rather than inferred from the thinking label.