GPT-6 Sol — Bid Bench
Bid Bench evaluates GPT-6 Sol in the context of heavy civil construction and estimating workflows.
The benchmark covers Documents, Takeoffs, Analysis, Knowledge, Apps and Predictions. Bid Bench 1.0 compares delivery quality across thinking settings, with forecasts and stopping behavior scored separately.
Strong delivery at a lower cost
Delivery scores weight Documents, Takeoffs, Analysis, Knowledge and Apps equally.
All four models shown. Time is measured on accepted work; scores include failures.
Lines connect each model’s best measured trade-offs. All measurements are available in the downloads.
GPT-6 Sol at Medium passes 94.4% of delivery attempts. The comparison settings pass 92.2% for Claude Fable 5.1 (Medium), 68.9% for Grok 4.7 (Fixed), 96.7% for GPT-5.6 Sol (Medium).
Swipe to compare all four models →
| GPT-6 SolMedium | Claude Fable 5.1Medium | Grok 4.7Fixed | GPT-5.6 SolMedium | |
|---|---|---|---|---|
| Cost / 100 accepted | $2.91 | — | — | $8.70 |
| Delivery score | 94.4% | 92.2% | 68.9% | 96.7% |
| DocumentsDocument reconciliation | 83.3% | 83.3% | 66.7% | 83.3% |
| TakeoffsQuantity calculations | 100.0% | 100.0% | 100.0% | 100.0% |
| AnalysisRecord analysis | 88.9% | 77.8% | 72.2% | 100.0% |
| KnowledgeConstruction concepts | 100.0% | 100.0% | 94.4% | 100.0% |
| AppsConstruction functions | 100.0% | 100.0% | 11.1% | 100.0% |
Medium for the three configurable models; Grok 4.7 has one fixed setting. These settings do not represent equal compute budgets. — means incomplete measurement. Full results: CSV · JSON.
Documents drive the remaining gap
The gains from thinking are concentrated in document reconciliation and record analysis.
Low passes 85/90 attempts, versus 75/90 with thinking Off. It also costs less per accepted result: $2.33 versus $3.08 per 100. Medium passes the same 85 attempts while using 13,383 tokens per attempt versus Low’s 12,034, and taking 10.4 seconds versus 7.4. High reaches 88/90; two service errors leave its total cost unknown.
Documents. Low passes 14/18, Medium 15/18 and High 16/18. Off, Low and Medium each stop on all three repeats of the largest reconciliation case; this includes transferring source data into the calculation tool.
Takeoffs. Every setting passes 18/18 schedule-based quantity and conversion checks. This result does not cover extracting geometry from drawings.
Analysis. Low and High pass 18/18, compared with Off’s 12/18. More thinking is not uniformly better: Medium passes 16/18.
Knowledge. Low and every higher setting pass 18/18, versus 16/18 with thinking Off.
Apps. Low passes 17/18; Medium and High pass 18/18. These tests exercise contractor functions, including procurement and scheduling.
Historical prices remain a stronger baseline
Both forecasts are scored against recorded outcomes. Zero error is perfect; beating the historical baseline means improving on a simple statistical forecast. The bidder baseline uses past district participation rates, with overall participation as a fallback; the price baseline uses historical median prices for the same item and unit.
Right of the dotted baseline is better. Bidder error is the Brier score × 100 (0–100); price error is a percentage. The two measures are not directly comparable.
The candidate panels cover 18 of 61 actual bidder/project appearances. Bidder error measures only that restricted panel.
High’s bidder error score is 6.73 out of 100 versus the historical baseline’s 6.77, a small difference on this sample. All settings perform worse on unit prices: even XHigh’s best mean group error is 60.0%, versus 35.7% for the historical baseline.
Harder reconciliation still causes failures
The tasks progress from routine work to the hardest cases. Higher lines mean more tasks completed correctly. Low passes all attempts through Expert, then 14/15 at Frontier and 11/15 at Stress. High reaches 13/15 at Stress.
15 attempts per difficulty level and thinking setting.
Every setting identifies all 12 deliberately blocked tasks correctly. Low’s three false stops show why that result must be read alongside willingness to finish solvable work.
| Thinking | Correct stops | False stops | Stop time |
|---|---|---|---|
| Off | 12/12 | 3/90 | 2.8 s |
| Low | 12/12 | 3/90 | 3.7 s |
| Medium | 12/12 | 3/90 | 4.7 s |
| High | 12/12 | 0/90 | 4.9 s |
| XHigh | 12/12 | 0/90 | 5.2 s |
| Max | 12/12 | 0/90 | 7.0 s |
False stops abandon solvable delivery tasks. Time is the median for correct stops, excluding timeouts.
Method
Each setting receives 30 delivery cases, repeated three times with frozen inputs, tools and acceptance checks. Cost includes failed attempts; tokens are input plus output per attempt; latency is the median for accepted results. Unknown complete cost or token measurements remain blank. These are controlled text, schedule and function tests, not scanned-plan takeoffs or the full Bidlo app. Method, uncertainty and source corrections · JSON · CSV.
Choosing a setting
Low is the strongest measured efficiency choice within Sol: 94.4% delivery at $2.33 per 100 accepted tasks. High improves completion to 97.8%, but incomplete cost data prevents a complete price comparison. Use the workflow results to decide whether that gain matters; Max did not improve on High here.