Claude Opus 5.5 on Bidlo Bench
Claude Opus 5.5 passed all 90 attempts on Bidlo Bench's 30 heavy civil estimating tasks, compared with 89 for Opus 5. The tasks cover supplied specification changes, bid-item extraction, quantities, delivered prices, and complete record proposals. The cost to run the model per successful result was $0.00387, versus $0.00659 for Opus 5.
Compare the models
Text v0.1 · 30 estimating tasks · 3 attempts each · September 26, 2026
Select a point or model row for its measured result and uncertainty.
Each point is one tested model configuration. The line joins observed best trade-offs; it is not a curve of untested reasoning settings.
| 1Adaptive thinking; requested effort: medium | 100.0%90/90 | $0.008890/90 billed | 47590/90 reported | 4.46 sn=90 passed |
| 1Adaptive thinking; requested effort: medium | 100.0%90/90 | $0.003990/90 billed | 49290/90 reported | 3.66 sn=90 passed |
| 1Reasoning effort: medium | 100.0%90/90 | $0.005390/90 billed | 30890/90 reported | 2.22 sn=90 passed |
| 1Reasoning effort: medium | 100.0%90/90 | $0.010690/90 billed | 30890/90 reported | 1.71 sn=90 passed |
| 1Reasoning effort: high | 100.0%90/90 | $0.001590/90 billed | 34890/90 reported | 2.25 sn=90 passed |
| 6Thinking budget: 8,000 tokens | 98.9%89/90 | $0.003790/90 billed | 98290/90 reported | 4.83 sn=89 passed |
| 6Adaptive thinking; requested effort: medium | 98.9%89/90 | $0.006590/90 billed | 55890/90 reported | 3.32 sn=89 passed |
| 6Requested thinking level: medium | 98.9%89/90 | $0.001890/90 billed | 93590/90 reported | 4.95 sn=89 passed |
| 6Reasoning effort: high | 98.9%89/90 | $0.002590/90 billed | 32690/90 reported | 2.16 sn=89 passed |
| 6Provider default; no explicit reasoning setting sent | 98.9%89/90 | $0.003390/90 billed | 1,63690/90 reported | 7.03 sn=89 passed |
| 11Requested thinking level: medium | 97.8%88/90 | $0.006890/90 billed | 78490/90 reported | 7.15 sn=88 passed |
| 11Requested thinking level: medium | 97.8%88/90 | Unknown88/90 billed | Unknown88/90 reported | 7.27 sn=88 passed |
| 13Adaptive thinking; requested effort: medium | 96.7%87/90 | $0.002790/90 billed | 56690/90 reported | 3.83 sn=87 passed |
| 13Reasoning effort: medium | 96.7%87/90 | $0.00016890/90 billed | 35090/90 reported | 2.45 sn=87 passed |
| 13Provider default; no explicit reasoning setting sent | 96.7%87/90 | $0.003290/90 billed | 1,93290/90 reported | 8.39 sn=87 passed |
| 16Requested reasoning effort: high | 95.6%86/90 | Unknown89/90 billed | Unknown89/90 reported | 2.17 sn=86 passed |
| 16Reasoning effort: medium | 95.6%86/90 | $0.001390/90 billed | 32290/90 reported | 2.04 sn=86 passed |
| 16Provider default; no explicit reasoning setting sent | 95.6%86/90 | Unknown88/90 billed | Unknown88/90 reported | 5.92 sn=86 passed |
| 19Reasoning effort: medium | 92.2%83/90 | $0.00007690/90 billed | 35490/90 reported | 2.02 sn=83 passed |
Equal scores share a rank. Small differences on 30 cases do not establish broad model superiority. Costs include failed attempts; incomplete billing or token totals remain unknown. Time is conditional on success. These scores cover supplied-text estimating tasks, not general intelligence or completed plan-to-workbook jobs.
One additional successful attempt
Opus 5's single unsuccessful answer omitted a fixed charge from an effective unit price, despite returning the correct quote total. The result illustrates why the benchmark checks related outputs independently: one correct number does not make the full comparison correct.
Opus 5.5 passed all three repeats of all 30 tasks. Opus 5 passed every repeat on 29 tasks. The new model's 90/90 result means no failures were observed in this finite test; it does not establish universal accuracy. The observed success-rate gain was 1.1 percentage points, with a 95% interval of 0.0–3.3 points.
Compliance with the requested machine-readable format was 87/90 for Opus 5.5 and 86/90 for Opus 5. Equivalent answer formats still passed the correctness checks.
Completion time and cost
Both models succeeded on all 30 tasks at least once. The median paired time ratio was 1.01×, with a 95% interval of 0.86–1.15×. The interval includes 1, so this test does not establish a speed advantage.
Billing covered every attempt. Opus 5.5's 90 attempts cost $0.34819, versus $0.58607 for Opus 5. Dividing all attempt costs by successful results gives $0.00387 and $0.00659 respectively. This includes Opus 5's unsuccessful attempt.
The clearest observed change here is lower inference cost while retaining all successful outcomes and resolving the older model's one miss. A larger set of difficult tasks is needed to establish how often that correctness difference recurs. We have yet to measure whether that result carries through to original plan review and finished Excel workbooks.
Method
Text v0.1 contains 30 synthetic tasks, repeated three times across 19 models: 1,710 attempts. Models received the same supplied text in fresh contexts, using recorded model settings. No original plan files were parsed, tools used, or Excel workbooks generated. The cases are not a representative sample of customer projects.
Grading v1.3 checks required outcomes and critical errors. After inspecting outputs, we revised grading to accept narrowly defined equivalent representations, applied uniformly to every model. Compliance with the requested answer format remains separate. Task definitions, arithmetic, graders, and results received independent AI review.
Elapsed time spans client request through response parsing. Paired ratios use task-level median successful times; 95% intervals use 2,000 source-group bootstrap resamples. The 180-second execution safeguard bounded requests. Errors retain observed time; timing creates no quality tiers.
Gateway inference cost is the charge to run the model, in USD. It includes unsuccessful attempts with known billing and is not a Bidlo customer charge. See aggregate results and settings, the 19-model comparison table (CSV), and the benchmark introduction.