GPT-6 Luna on Bidlo Bench
GPT-6 Luna passed 83 of 90 attempts on Bidlo Bench's 30 heavy civil estimating tasks, compared with 87 for GPT-5.6 Luna. The tasks cover supplied specification changes, bid-item extraction, quantities, delivered prices, and complete record proposals. It cost less to run and used less elapsed time on shared successful tasks, with more incorrect outcomes.
Compare the models
Text v0.1 · 30 estimating tasks · 3 attempts each · September 26, 2026
Select a point or model row for its measured result and uncertainty.
Each point is one tested model configuration. The line joins observed best trade-offs; it is not a curve of untested reasoning settings.
| 1Adaptive thinking; requested effort: medium | 100.0%90/90 | $0.008890/90 billed | 47590/90 reported | 4.46 sn=90 passed |
| 1Adaptive thinking; requested effort: medium | 100.0%90/90 | $0.003990/90 billed | 49290/90 reported | 3.66 sn=90 passed |
| 1Reasoning effort: medium | 100.0%90/90 | $0.005390/90 billed | 30890/90 reported | 2.22 sn=90 passed |
| 1Reasoning effort: medium | 100.0%90/90 | $0.010690/90 billed | 30890/90 reported | 1.71 sn=90 passed |
| 1Reasoning effort: high | 100.0%90/90 | $0.001590/90 billed | 34890/90 reported | 2.25 sn=90 passed |
| 6Thinking budget: 8,000 tokens | 98.9%89/90 | $0.003790/90 billed | 98290/90 reported | 4.83 sn=89 passed |
| 6Adaptive thinking; requested effort: medium | 98.9%89/90 | $0.006590/90 billed | 55890/90 reported | 3.32 sn=89 passed |
| 6Requested thinking level: medium | 98.9%89/90 | $0.001890/90 billed | 93590/90 reported | 4.95 sn=89 passed |
| 6Reasoning effort: high | 98.9%89/90 | $0.002590/90 billed | 32690/90 reported | 2.16 sn=89 passed |
| 6Provider default; no explicit reasoning setting sent | 98.9%89/90 | $0.003390/90 billed | 1,63690/90 reported | 7.03 sn=89 passed |
| 11Requested thinking level: medium | 97.8%88/90 | $0.006890/90 billed | 78490/90 reported | 7.15 sn=88 passed |
| 11Requested thinking level: medium | 97.8%88/90 | Unknown88/90 billed | Unknown88/90 reported | 7.27 sn=88 passed |
| 13Adaptive thinking; requested effort: medium | 96.7%87/90 | $0.002790/90 billed | 56690/90 reported | 3.83 sn=87 passed |
| 13Reasoning effort: medium | 96.7%87/90 | $0.00016890/90 billed | 35090/90 reported | 2.45 sn=87 passed |
| 13Provider default; no explicit reasoning setting sent | 96.7%87/90 | $0.003290/90 billed | 1,93290/90 reported | 8.39 sn=87 passed |
| 16Requested reasoning effort: high | 95.6%86/90 | Unknown89/90 billed | Unknown89/90 reported | 2.17 sn=86 passed |
| 16Reasoning effort: medium | 95.6%86/90 | $0.001390/90 billed | 32290/90 reported | 2.04 sn=86 passed |
| 16Provider default; no explicit reasoning setting sent | 95.6%86/90 | Unknown88/90 billed | Unknown88/90 reported | 5.92 sn=86 passed |
| 19Reasoning effort: medium | 92.2%83/90 | $0.00007690/90 billed | 35490/90 reported | 2.02 sn=83 passed |
Equal scores share a rank. Small differences on 30 cases do not establish broad model superiority. Costs include failed attempts; incomplete billing or token totals remain unknown. Time is conditional on success. These scores cover supplied-text estimating tasks, not general intelligence or completed plan-to-workbook jobs.
Where Luna lost successful outcomes
Both models omitted a fixed charge from an effective unit price in three attempts. GPT-6 Luna had four additional failures: two incorrect sums of selected bid items, one incomplete returned record set, and one incorrect delivered-quote calculation despite identifying the right supplier.
These are separate checks. Choosing the right items or supplier does not guarantee the arithmetic is correct. Likewise, correctly proposing one requested record update does not satisfy a request to return the full record set when another required record is missing. This was a text proposal, not a live database operation.
GPT-6 Luna passed all three repeats on 26 of 30 tasks, versus 29 for GPT-5.6 Luna. Compliance with the requested machine-readable format was 80/90 and 83/90 respectively. The observed success-rate difference was −4.4 percentage points, with a 95% interval of −10.0 to 0.0 points.
Completion time and cost
The paired timing comparison covers 29 tasks where both models succeeded at least once. The median time ratio was 0.87×, with a 95% interval of 0.81–0.95×. The interval falls below 1, consistent with less elapsed time on successful attempts in this shared subset. The task excluded from this comparison remains in the correctness results; otherwise a fast model could look better simply by failing harder work.
Billing covered all attempts: $0.006855 for GPT-6 Luna and $0.015127 for GPT-5.6 Luna. Including failures, cost per successful result was $0.0000826 and $0.0001739 respectively.
GPT-6 Luna's observed tradeoff is clear: lower inference cost with fewer successful outcomes. That can matter for routine work with independent checks, but these results do not justify relaxing review of quote arithmetic, totals, or returned records. The next useful test is whether those additional failure patterns persist on a larger task set.
Method
Text v0.1 contains 30 synthetic tasks, repeated three times across 19 models: 1,710 attempts. Models received the same supplied text in fresh contexts, using recorded model settings. No original plan files were parsed, tools used, or Excel workbooks generated. The cases are not a representative sample of customer projects.
Grading v1.3 checks required outcomes and critical errors. After inspecting outputs, we revised grading to accept narrowly defined equivalent representations, applied uniformly to every model. Compliance with the requested answer format remains separate. Task definitions, arithmetic, graders, and results received independent AI review.
Elapsed time spans client request through response parsing. Paired ratios use task-level median successful times; 95% intervals use 2,000 source-group bootstrap resamples. The 180-second execution safeguard bounded requests. Errors retain observed time; timing creates no quality tiers.
Gateway inference cost is the charge to run the model, in USD. It includes unsuccessful attempts with known billing and is not a Bidlo customer charge. See aggregate results and settings, the 19-model comparison table (CSV), and the benchmark introduction.