Ideas

GPT-6 Luna — Bid Bench

Matt Wolfe3 min read

Bid Bench evaluates GPT-6 Luna in the context of heavy civil construction and estimating workflows.

The benchmark covers Documents, Takeoffs, Analysis, Knowledge, Apps and Predictions. Bid Bench 1.0 compares delivery quality across thinking settings, with forecasts and stopping behavior scored separately.

Low cost comes with more failures

Delivery scores weight the first five workflows equally.

Bid Bench delivery score20%30%40%50%60%70%80%90%100%0 s10 s20 s30 sGPT-6 Luna · Max: 92.2% passed; 18.2 sGPT-6 Luna · XHigh: 88.9% passed; 11.5 sGPT-6 Luna · Medium: 78.9% passed; 9.6 sGPT-6 Luna · Low: 70.0% passed; 6.9 sClaude Haiku 4.5 · High: 37.8% passed; 26.8 sClaude Haiku 4.5 · Off: 31.1% passed; 25.1 sGemini 3.8 Flash · Low: 83.3% passed; 23.1 sGPT-5.6 Luna · XHigh: 81.1% passed; 13.6 sGPT-5.6 Luna · Low: 65.6% passed; 6.4 sGPT-5.6 Luna · Off: 56.7% passed; 4.6 sGPT-6 LunaGemini 3.8 FlashGPT-5.6 LunaClaude Haiku 4.5Median accepted completion time · seconds

All four models shown. Time is measured on accepted work; scores include failures.

Lines connect each model’s best measured trade-offs. All measurements are available in the downloads.

At Medium, Luna passes 78.9% of delivery attempts, versus 73.3% for GPT-5.6 Luna, 57.8% for Gemini 3.8 Flash and 32.2% for Claude Haiku 4.5. Luna improves on its predecessor in Documents, Takeoffs and Knowledge; the older model still leads in Analysis and Apps.

Swipe to compare all four models →

GPT-6 LunaMediumClaude Haiku 4.5MediumGemini 3.8 FlashMediumGPT-5.6 LunaMedium
Cost / 100 accepted——$0.46
Delivery score32.2%57.8%73.3%
DocumentsDocument reconciliation50.0%61.1%66.7%
TakeoffsQuantity calculations33.3%88.9%50.0%
AnalysisRecord analysis16.7%38.9%66.7%
KnowledgeConstruction concepts55.6%100.0%88.9%
AppsConstruction functions5.6%0.0%94.4%

Medium for all four models; providers do not use identical compute budgets. — means incomplete measurement. Full results: CSV · JSON.


Thinking helps some workflows more

Knowledge improves early; reliable record analysis needs more thinking in this set.

Low passes 63/90 attempts at $0.17 per 100 accepted tasks. Medium improves to 71/90 at $0.22. High also passes 71/90, with slightly more tokens and time. Max reaches 83/90, but incomplete telemetry prevents a complete cost comparison. Median accepted latency rises from 6.9 seconds at Low to 18.2 at Max.

GPT-6 Luna across workflowsOffLowMediumHighXHighMax20%30%40%50%60%70%80%90%100%Documents44.466.772.261.177.883.3Takeoffs83.388.983.377.888.988.9Analysis33.350.050.061.183.3100.0Knowledge38.955.6100.0100.0100.0100.0Apps83.388.988.994.494.488.9Bid Bench 1.0 · measured September 28, 2026

Documents. Low passes 12/18 and Medium 13/18. Max reaches 15/18, leaving document reconciliation short of full completion.

Takeoffs. Low and Max both pass 16/18. An inspected Off response divides by 27 twice when converting cubic feet to cubic yards; additional thinking does not eliminate every quantity error.

Analysis. Low and Medium pass 9/18; Max reaches 18/18. An inspected Medium answer selects the wrong eligible winning bid.

Knowledge. Medium through Max pass 18/18, compared with Low’s 10/18. This is the clearest early gain from thinking.

Apps. Low and Medium pass 16/18; High and XHigh reach 17/18. Max returns to 16/18, so the highest setting is not uniformly strongest.


Forecasts do not beat the baselines

Both forecasts are scored against recorded outcomes. Zero error is perfect; beating the historical baseline means improving on a simple statistical forecast. The bidder baseline uses past district participation rates, with overall participation as a fallback; the price baseline uses historical median prices for the same item and unit.

Bidder probabilitiesBaseline 6.77LowGPT-6 Luna · Low: 9.59 error; historical baseline 6.77; farther right is better9.59MediumGPT-6 Luna · Medium: 6.89 error; historical baseline 6.77; farther right is better6.89HighGPT-6 Luna · High: 6.80 error; historical baseline 6.77; farther right is better6.80XHighGPT-6 Luna · XHigh: 6.81 error; historical baseline 6.77; farther right is better6.81MaxGPT-6 Luna · Max: 6.87 error; historical baseline 6.77; farther right is better6.87ErrorPerfect (0)2.557.510Off omitted: 8/12 valid batchesUnit pricesBaseline 35.7%LowGPT-6 Luna · Low: 41.8% error; historical baseline 35.7%; farther right is better41.8%MediumGPT-6 Luna · Medium: 44.6% error; historical baseline 35.7%; farther right is better44.6%HighGPT-6 Luna · High: 58.8% error; historical baseline 35.7%; farther right is better58.8%XHighGPT-6 Luna · XHigh: 52.9% error; historical baseline 35.7%; farther right is better52.9%MaxGPT-6 Luna · Max: 54.3% error; historical baseline 35.7%; farther right is better54.3%ErrorPerfect (0)20%40%60%Off omitted: 10/12 valid batches

Right of the dotted baseline is better. Bidder error is the Brier score × 100 (0–100); price error is a percentage. The two measures are not directly comparable.

The candidate panels cover 18 of 61 actual bidder/project appearances. Bidder error measures only that restricted panel.

Among complete settings, High has the lowest bidder error score, 6.80 out of 100, versus the baseline’s 6.77. Low has the lowest price error, 41.8%, versus 35.7%. Thinking Off supplies only 8/12 valid bidder batches and 10/12 price batches, so it has no complete forecast point. These retrospective results do not validate prospective prediction accuracy.


Hard cases benefit from higher settings

We test construction tasks at six difficulty levels, from routine work to the hardest cases. Each line shows how often the model completes them correctly at a different thinking setting. On the hardest tasks, Max passes 12 of 15 attempts; Off passes 1.

Tasks completed correctly0%25%50%75%100%123456Off · difficulty 1 (Routine): 13/15 completed correctly (86.7%)Off · difficulty 2 (Multi-step): 13/15 completed correctly (86.7%)Off · difficulty 3 (Complex): 9/15 completed correctly (60.0%)Off · difficulty 4 (Expert): 9/15 completed correctly (60.0%)Off · difficulty 5 (Frontier): 6/15 completed correctly (40.0%)Off · difficulty 6 (Stress): 1/15 completed correctly (6.7%)Low · difficulty 1 (Routine): 15/15 completed correctly (100.0%)Low · difficulty 2 (Multi-step): 12/15 completed correctly (80.0%)Low · difficulty 3 (Complex): 11/15 completed correctly (73.3%)Low · difficulty 4 (Expert): 11/15 completed correctly (73.3%)Low · difficulty 5 (Frontier): 9/15 completed correctly (60.0%)Low · difficulty 6 (Stress): 5/15 completed correctly (33.3%)Medium · difficulty 1 (Routine): 14/15 completed correctly (93.3%)Medium · difficulty 2 (Multi-step): 14/15 completed correctly (93.3%)Medium · difficulty 3 (Complex): 13/15 completed correctly (86.7%)Medium · difficulty 4 (Expert): 14/15 completed correctly (93.3%)Medium · difficulty 5 (Frontier): 9/15 completed correctly (60.0%)Medium · difficulty 6 (Stress): 7/15 completed correctly (46.7%)High · difficulty 1 (Routine): 12/15 completed correctly (80.0%)High · difficulty 2 (Multi-step): 14/15 completed correctly (93.3%)High · difficulty 3 (Complex): 13/15 completed correctly (86.7%)High · difficulty 4 (Expert): 12/15 completed correctly (80.0%)High · difficulty 5 (Frontier): 12/15 completed correctly (80.0%)High · difficulty 6 (Stress): 8/15 completed correctly (53.3%)XHigh · difficulty 1 (Routine): 13/15 completed correctly (86.7%)XHigh · difficulty 2 (Multi-step): 15/15 completed correctly (100.0%)XHigh · difficulty 3 (Complex): 15/15 completed correctly (100.0%)XHigh · difficulty 4 (Expert): 13/15 completed correctly (86.7%)XHigh · difficulty 5 (Frontier): 12/15 completed correctly (80.0%)XHigh · difficulty 6 (Stress): 12/15 completed correctly (80.0%)Max · difficulty 1 (Routine): 13/15 completed correctly (86.7%)Max · difficulty 2 (Multi-step): 15/15 completed correctly (100.0%)Max · difficulty 3 (Complex): 14/15 completed correctly (93.3%)Max · difficulty 4 (Expert): 14/15 completed correctly (93.3%)Max · difficulty 5 (Frontier): 15/15 completed correctly (100.0%)Max · difficulty 6 (Stress): 12/15 completed correctly (80.0%)XHigh 80.0%Max 80.0%High 53.3%Medium 46.7%Low 33.3%Off 6.7%Task difficulty · 1 easiest, 6 hardest

15 attempts per difficulty level and thinking setting.

High through Max identify all 12 blocked tasks correctly. False stops fall from 12/90 with thinking Off to zero at Max. The trade-off is longer time to a correct stop.

ThinkingCorrect stopsFalse stopsStop time
Off9/1212/901.7 s
Low11/1210/903.6 s
Medium10/123/903.5 s
High12/123/903.7 s
XHigh12/121/904.6 s
Max12/120/906.1 s

False stops abandon solvable delivery tasks. Time is the median for correct stops, excluding timeouts.


Method

Each setting receives 30 delivery cases, repeated three times with frozen inputs, tools and acceptance checks. Cost includes failed attempts; tokens are input plus output per attempt; latency is the median for accepted results. Unknown complete cost or token measurements remain blank. These are controlled text, schedule and function tests, not scanned-plan takeoffs or the full Bidlo app. Method, uncertainty and source corrections · JSON · CSV.


Choosing a setting

Luna fits work where low cost and explicit verification matter. Medium improves Knowledge substantially over Low; Max is stronger for Analysis.

Filed under: Ideas

Author:
Matt Wolfe

Explore the letting market with Bid.

Choose what you get.

Pick what gets sent.

  • Product
  • Research
  • Stories