Ideas

GPT-6 Sol — Bid Bench

Matt Wolfe3 min de lectura

Bid Bench evaluates GPT-6 Sol in the context of heavy civil construction and estimating workflows.

The benchmark covers Documents, Takeoffs, Analysis, Knowledge, Apps and Predictions. Bid Bench 1.0 compares delivery quality across thinking settings, with forecasts and stopping behavior scored separately.

Strong delivery at a lower cost

Delivery scores weight Documents, Takeoffs, Analysis, Knowledge and Apps equally.

Bid Bench delivery score60%70%80%90%100%0 s10 s20 s30 s40 s50 sGPT-6 Sol · High: 97.8% passed; 11.7 sGPT-6 Sol · Low: 94.4% passed; 7.4 sClaude Fable 5.1 · High: 93.3% passed; 32.8 sClaude Fable 5.1 · Low: 90.0% passed; 31.5 sGrok 4.7 · Fixed: 68.9% passed; 48.8 sGPT-5.6 Sol · XHigh: 100.0% passed; 16.8 sGPT-5.6 Sol · Medium: 96.7% passed; 14.4 sGPT-5.6 Sol · Low: 95.6% passed; 13.1 sGPT-5.6 Sol · Off: 84.4% passed; 9.1 sGPT-6 SolGPT-5.6 SolClaude Fable 5.1Grok 4.7Median accepted completion time · seconds

All four models shown. Time is measured on accepted work; scores include failures.

Lines connect each model’s best measured trade-offs. All measurements are available in the downloads.

GPT-6 Sol at Medium passes 94.4% of delivery attempts. The comparison settings pass 92.2% for Claude Fable 5.1 (Medium), 68.9% for Grok 4.7 (Fixed), 96.7% for GPT-5.6 Sol (Medium).

Swipe to compare all four models →

GPT-6 SolMediumClaude Fable 5.1MediumGrok 4.7FixedGPT-5.6 SolMedium
Cost / 100 accepted——$8.70
Delivery score92.2%68.9%96.7%
DocumentsDocument reconciliation83.3%66.7%83.3%
TakeoffsQuantity calculations100.0%100.0%100.0%
AnalysisRecord analysis77.8%72.2%100.0%
KnowledgeConstruction concepts100.0%94.4%100.0%
AppsConstruction functions100.0%11.1%100.0%

Medium for the three configurable models; Grok 4.7 has one fixed setting. These settings do not represent equal compute budgets. — means incomplete measurement. Full results: CSV · JSON.


Documents drive the remaining gap

The gains from thinking are concentrated in document reconciliation and record analysis.

Low passes 85/90 attempts, versus 75/90 with thinking Off. It also costs less per accepted result: $2.33 versus $3.08 per 100. Medium passes the same 85 attempts while using 13,383 tokens per attempt versus Low’s 12,034, and taking 10.4 seconds versus 7.4. High reaches 88/90; two service errors leave its total cost unknown.

GPT-6 Sol across workflowsOffLowMediumHighXHighMax50%60%70%80%90%100%Documents61.177.883.388.983.383.3Takeoffs100.0100.0100.0100.0100.0100.0Analysis66.7100.088.9100.094.4100.0Knowledge88.9100.0100.0100.0100.0100.0Apps100.094.4100.0100.0100.094.4Bid Bench 1.0 · measured September 28, 2026

Documents. Low passes 14/18, Medium 15/18 and High 16/18. Off, Low and Medium each stop on all three repeats of the largest reconciliation case; this includes transferring source data into the calculation tool.

Takeoffs. Every setting passes 18/18 schedule-based quantity and conversion checks. This result does not cover extracting geometry from drawings.

Analysis. Low and High pass 18/18, compared with Off’s 12/18. More thinking is not uniformly better: Medium passes 16/18.

Knowledge. Low and every higher setting pass 18/18, versus 16/18 with thinking Off.

Apps. Low passes 17/18; Medium and High pass 18/18. These tests exercise contractor functions, including procurement and scheduling.


Historical prices remain a stronger baseline

Both forecasts are scored against recorded outcomes. Zero error is perfect; beating the historical baseline means improving on a simple statistical forecast. The bidder baseline uses past district participation rates, with overall participation as a fallback; the price baseline uses historical median prices for the same item and unit.

Bidder probabilitiesBaseline 6.77OffGPT-6 Sol · Off: 7.07 error; historical baseline 6.77; farther right is better7.07LowGPT-6 Sol · Low: 7.10 error; historical baseline 6.77; farther right is better7.10MediumGPT-6 Sol · Medium: 6.82 error; historical baseline 6.77; farther right is better6.82HighGPT-6 Sol · High: 6.73 error; historical baseline 6.77; farther right is better6.73XHighGPT-6 Sol · XHigh: 6.78 error; historical baseline 6.77; farther right is better6.78MaxGPT-6 Sol · Max: 6.75 error; historical baseline 6.77; farther right is better6.75ErrorPerfect (0)2.557.5Unit pricesBaseline 35.7%OffGPT-6 Sol · Off: 62.8% error; historical baseline 35.7%; farther right is better62.8%LowGPT-6 Sol · Low: 60.6% error; historical baseline 35.7%; farther right is better60.6%MediumGPT-6 Sol · Medium: 63.1% error; historical baseline 35.7%; farther right is better63.1%HighGPT-6 Sol · High: 61.7% error; historical baseline 35.7%; farther right is better61.7%XHighGPT-6 Sol · XHigh: 60.0% error; historical baseline 35.7%; farther right is better60.0%MaxGPT-6 Sol · Max: 61.3% error; historical baseline 35.7%; farther right is better61.3%ErrorPerfect (0)20%40%60%80%

Right of the dotted baseline is better. Bidder error is the Brier score × 100 (0–100); price error is a percentage. The two measures are not directly comparable.

The candidate panels cover 18 of 61 actual bidder/project appearances. Bidder error measures only that restricted panel.

High’s bidder error score is 6.73 out of 100 versus the historical baseline’s 6.77, a small difference on this sample. All settings perform worse on unit prices: even XHigh’s best mean group error is 60.0%, versus 35.7% for the historical baseline.


Harder reconciliation still causes failures

The tasks progress from routine work to the hardest cases. Higher lines mean more tasks completed correctly. Low passes all attempts through Expert, then 14/15 at Frontier and 11/15 at Stress. High reaches 13/15 at Stress.

Tasks completed correctly0%25%50%75%100%123456Off · difficulty 1 (Routine): 15/15 completed correctly (100.0%)Off · difficulty 2 (Multi-step): 14/15 completed correctly (93.3%)Off · difficulty 3 (Complex): 14/15 completed correctly (93.3%)Off · difficulty 4 (Expert): 13/15 completed correctly (86.7%)Off · difficulty 5 (Frontier): 12/15 completed correctly (80.0%)Off · difficulty 6 (Stress): 7/15 completed correctly (46.7%)Low · difficulty 1 (Routine): 15/15 completed correctly (100.0%)Low · difficulty 2 (Multi-step): 15/15 completed correctly (100.0%)Low · difficulty 3 (Complex): 15/15 completed correctly (100.0%)Low · difficulty 4 (Expert): 15/15 completed correctly (100.0%)Low · difficulty 5 (Frontier): 14/15 completed correctly (93.3%)Low · difficulty 6 (Stress): 11/15 completed correctly (73.3%)Medium · difficulty 1 (Routine): 15/15 completed correctly (100.0%)Medium · difficulty 2 (Multi-step): 15/15 completed correctly (100.0%)Medium · difficulty 3 (Complex): 15/15 completed correctly (100.0%)Medium · difficulty 4 (Expert): 15/15 completed correctly (100.0%)Medium · difficulty 5 (Frontier): 15/15 completed correctly (100.0%)Medium · difficulty 6 (Stress): 10/15 completed correctly (66.7%)High · difficulty 1 (Routine): 15/15 completed correctly (100.0%)High · difficulty 2 (Multi-step): 15/15 completed correctly (100.0%)High · difficulty 3 (Complex): 15/15 completed correctly (100.0%)High · difficulty 4 (Expert): 15/15 completed correctly (100.0%)High · difficulty 5 (Frontier): 15/15 completed correctly (100.0%)High · difficulty 6 (Stress): 13/15 completed correctly (86.7%)XHigh · difficulty 1 (Routine): 15/15 completed correctly (100.0%)XHigh · difficulty 2 (Multi-step): 15/15 completed correctly (100.0%)XHigh · difficulty 3 (Complex): 15/15 completed correctly (100.0%)XHigh · difficulty 4 (Expert): 15/15 completed correctly (100.0%)XHigh · difficulty 5 (Frontier): 15/15 completed correctly (100.0%)XHigh · difficulty 6 (Stress): 11/15 completed correctly (73.3%)Max · difficulty 1 (Routine): 15/15 completed correctly (100.0%)Max · difficulty 2 (Multi-step): 14/15 completed correctly (93.3%)Max · difficulty 3 (Complex): 15/15 completed correctly (100.0%)Max · difficulty 4 (Expert): 15/15 completed correctly (100.0%)Max · difficulty 5 (Frontier): 15/15 completed correctly (100.0%)Max · difficulty 6 (Stress): 12/15 completed correctly (80.0%)High 86.7%Max 80.0%Low 73.3%XHigh 73.3%Medium 66.7%Off 46.7%Task difficulty · 1 easiest, 6 hardest

15 attempts per difficulty level and thinking setting.

Every setting identifies all 12 deliberately blocked tasks correctly. Low’s three false stops show why that result must be read alongside willingness to finish solvable work.

ThinkingCorrect stopsFalse stopsStop time
Off12/123/902.8 s
Low12/123/903.7 s
Medium12/123/904.7 s
High12/120/904.9 s
XHigh12/120/905.2 s
Max12/120/907.0 s

False stops abandon solvable delivery tasks. Time is the median for correct stops, excluding timeouts.


Method

Each setting receives 30 delivery cases, repeated three times with frozen inputs, tools and acceptance checks. Cost includes failed attempts; tokens are input plus output per attempt; latency is the median for accepted results. Unknown complete cost or token measurements remain blank. These are controlled text, schedule and function tests, not scanned-plan takeoffs or the full Bidlo app. Method, uncertainty and source corrections · JSON · CSV.


Choosing a setting

Low is the strongest measured efficiency choice within Sol: 94.4% delivery at $2.33 per 100 accepted tasks. High improves completion to 97.8%, but incomplete cost data prevents a complete price comparison. Use the workflow results to decide whether that gain matters; Max did not improve on High here.

Archivado en: Ideas

Autor:
Matt Wolfe

Explora el mercado de licitaciones con Licitar.

Elige lo que recibes.

Elige lo que te llega.

  • Producto
  • Investigación
  • Historias