GPT-6 Sol on Bid Bench
Bid Bench evaluates GPT-6 Sol in the context of heavy civil construction and estimating workflows.
The benchmark covers Documents, Takeoffs, Analysis, Knowledge, Apps and Predictions. Bid Bench 1.0 compares delivery quality across thinking settings, with forecasts and stopping behavior scored separately.
Low captures most of the gain
Delivery scores weight Documents, Takeoffs, Analysis, Knowledge and Apps equally.
Measured thinking settings and selected alternatives · Logarithmic horizontal axis · 3 missing measurements retained in the table
| Model · Thinking | Pass | Cost¹ | Tokens² | Latency³ |
|---|---|---|---|---|
| 83.3% | $3.08 | 17,675 | 7.9 s | |
| 94.4% | $2.33 | 12,034 | 7.4 s | |
| 94.4% | $2.91 | 13,383 | 10.4 s | |
| 97.8% | — | — | 11.7 s | |
| 95.6% | — | — | 13.5 s | |
| 95.6% | — | — | 15.8 s | |
| 70.0% | $0.17 | 8,450 | 6.9 s | |
| 100.0% | $9.83 | 21,906 | 17.6 s | |
| 94.4% | $4.93 | 19,705 | 10.9 s |
All 75 configurations · 18 models
| Model · Thinking | Pass | Cost¹ | Tokens² | Latency³ |
|---|---|---|---|---|
| 100.0% | $9.83 | 21,906 | 17.6 s | |
| 100.0% | $11.66 | 26,419 | 16.8 s | |
| 100.0% | $11.84 | 11,621 | 10.7 s | |
| 100.0% | $13.96 | 12,848 | 12.8 s | |
| 100.0% | $15.15 | 25,258 | 21.9 s | |
| 100.0% | $15.51 | 26,053 | 23.4 s | |
| 100.0% | $17.42 | 29,689 | 23.3 s | |
| 100.0% | $17.87 | 13,921 | 18.9 s | |
| 98.9% | $14.28 | 23,131 | 22.8 s | |
| 98.9% | — | — | 22.2 s | |
| 98.9% | — | — | 27.7 s | |
| 97.8% | — | — | 11.7 s | |
| 96.7% | $8.70 | 21,280 | 14.4 s | |
| 96.7% | — | — | 17.1 s | |
| 95.6% | $6.24 | 16,797 | 13.1 s | |
| 95.6% † | — | — | 31.3 s | |
| 95.6% | — | — | 15.8 s | |
| 95.6% | — | — | 13.5 s | |
| 94.4% | $2.33 | 12,034 | 7.4 s | |
| 94.4% | $2.91 | 13,383 | 10.4 s | |
| 94.4% | $4.93 | 19,705 | 10.9 s | |
| 94.4% | — | — | 32.8 s | |
| 94.4% | — | — | 31.5 s | |
| 93.3% | — | — | 32.8 s | |
| 93.3% † | — | — | 31.2 s | |
| 93.3% | — | — | 30.3 s | |
| 92.2% | — | — | 35.5 s | |
| 92.2% | — | — | 33.5 s | |
| 92.2% † | — | — | 29.7 s | |
| 92.2% | — | — | 19.3 s | |
| 92.2% | — | — | 18.2 s | |
| 91.1% | — | — | 35.0 s | |
| 90.0% | — | — | 31.5 s | |
| 90.0% | — | — | 24.5 s | |
| 88.9% | — | — | 35.6 s | |
| 88.9% | — | — | 11.5 s | |
| 87.8% | — | — | 11.5 s | |
| 86.7% | $4.34 | 17,737 | 9.4 s | |
| 85.6% | $3.77 | 15,669 | 8.6 s | |
| 85.6% | — | — | 49.3 s | |
| 84.4% | $7.84 | 19,566 | 9.1 s | |
| 83.3% | $3.08 | 17,675 | 7.9 s | |
| 83.3% † | — | — | 23.1 s | |
| 81.1% | — | — | 51.7 s | |
| 81.1% | — | — | 39.1 s | |
| 81.1% | — | — | 13.6 s | |
| 78.9% | $0.22 | 15,802 | 9.6 s | |
| 78.9% | $0.25 | 15,844 | 10.0 s | |
| 78.9% | $0.84 | 27,834 | 16.0 s | |
| 78.9% | — | — | 41.3 s | |
| 78.9% | — | — | 66.6 s | |
| 78.9% | — | — | 27.3 s | |
| 76.7% † | — | — | 42.8 s | |
| 75.6% † | — | — | 38.8 s | |
| 74.4% | — | — | 53.0 s | |
| 73.3% | $0.46 | 12,354 | 13.9 s | |
| 70.0% | $0.17 | 8,450 | 6.9 s | |
| 68.9% † | — | — | 48.8 s | |
| 67.8% | — | — | 67.3 s | |
| 65.6% | $5.01 | 13,514 | 4.7 s | |
| 65.6% | — | — | 6.4 s | |
| 63.3% | — | — | 26.7 s | |
| 63.3% | — | — | 19.7 s | |
| 57.8% | $3.22 | 47,659 | 13.3 s | |
| 57.8% | — | — | 36.8 s | |
| 56.7% | $0.26 | 20,711 | 8.0 s | |
| 56.7% | $0.57 | 13,523 | 4.6 s | |
| 47.8% | — | — | 78.1 s | |
| 46.7% | $4.52 | 47,105 | 13.5 s | |
| 46.7% | $6.49 | 66,166 | 30.4 s | |
| 44.4% † | — | — | 46.2 s | |
| 37.8% † | — | — | 26.8 s | |
| 34.4% † | — | — | 11.1 s | |
| 32.2% † | — | — | 29.3 s | |
| 31.1% † | — | — | 25.1 s |
¹ USD per 100 accepted delivery tasks, including failed attempts. ² Mean input + output tokens per attempt. ³ Median seconds on accepted tasks. Missing telemetry is shown as —. Forecasts and judgment are scored separately.
† Includes attempts interrupted by our test host. These do not establish a reasoning error.
GPT-6 Sol · Medium: 85/90 passed. Case-resampling interval 86.7–100.0%. $2.91 per 100 accepted; 10.4 s median on accepted tasks.
Results JSON · CSV · Method · September 28, 2026 UTC
Low passes 85/90 attempts, versus 75/90 with thinking Off. It also costs less per accepted result: $2.33 versus $3.08 per 100. Medium passes the same 85 attempts while using 13,383 tokens per attempt versus Low’s 12,034, and taking 10.4 seconds versus 7.4. High reaches 88/90; two service errors leave its total cost unknown.
Documents drive the remaining gap
The gains from thinking are concentrated in document reconciliation and record analysis.
Documents. Low passes 14/18, Medium 15/18 and High 16/18. Off, Low and Medium each stop on all three repeats of the largest reconciliation case; this includes transferring source data into the calculation tool.
Takeoffs. Every setting passes 18/18 schedule-based quantity and conversion checks. This result does not cover extracting geometry from drawings.
Analysis. Low and High pass 18/18, compared with Off’s 12/18. More thinking is not uniformly better: Medium passes 16/18.
Knowledge. Low and every higher setting pass 18/18, versus 16/18 with thinking Off.
Apps. Low passes 17/18; Medium and High pass 18/18. These tests exercise contractor functions, including procurement and scheduling.
Historical prices remain a stronger baseline
Forecasts use a small retrospective Texas cohort, evaluated separately from delivery.
The candidate panels cover 18 of 61 actual bidder/project appearances. Bidder error measures only that restricted panel.
High’s bidder Brier score is 0.0673 versus the historical baseline’s 0.0677, a small difference on this sample. All settings perform worse on unit prices: even XHigh’s best mean group error is 60.0%, versus 35.7% for the historical baseline.
Harder reconciliation still causes failures
Low passes all attempts through Expert, then 14/15 at Frontier and 11/15 at Stress. High reaches 13/15 at Stress.
Every setting identifies all 12 deliberately blocked tasks correctly. Low’s three false stops show why that result must be read alongside willingness to finish solvable work.
| Thinking | Correct stops | False stops | Stop time |
|---|---|---|---|
| Off | 12/12 | 3/90 | 2.8 s |
| Low | 12/12 | 3/90 | 3.7 s |
| Medium | 12/12 | 3/90 | 4.7 s |
| High | 12/12 | 0/90 | 4.9 s |
| XHigh | 12/12 | 0/90 | 5.2 s |
| Max | 12/12 | 0/90 | 7.0 s |
False stops abandon solvable delivery tasks. Time is the median for correct stops, excluding timeouts.
Method
Each setting receives 30 delivery cases, repeated three times with frozen inputs, tools and acceptance checks. Cost includes failed attempts; tokens are input plus output per attempt; latency is the median for accepted results. Unknown complete cost or token measurements remain blank. These are controlled text, schedule and function tests, not scanned-plan takeoffs or the full Bidlo app. Method, uncertainty and source corrections · JSON · CSV.
Choosing a setting
Low is the strongest measured efficiency choice within Sol: 94.4% delivery at $2.33 per 100 accepted tasks. High improves completion to 97.8%, but incomplete cost data prevents a complete price comparison. Use the workflow results to decide whether that gain matters; Max did not improve on High here.