Introducing Bidlo Bench
Updated September 26, 2026.
Bidlo Bench measures how well AI handles heavy civil estimating tasks. We built it to answer practical questions: whether a model can follow an addendum, extract the right bid items, check a quantity, compare delivered prices, and keep every required record in a report.
A useful answer has to survive those checks. A quote total can be right while its unit price leaves out mobilization. A model can select the right items and sum them incorrectly. A revised specification can change the answer without changing the wording of the estimator's request.
The first edition tests these decisions using supplied text. Reading original plan sheets and producing finished Excel workbooks are part of the expansion we are building toward; this edition did not execute those workflows.
Start with the work
Each task has a defined request, supporting information, and checks for the result. The current set covers specification and revision interpretation, bid-item extraction, quantity and haul calculations, quote review, and project-record completeness. When necessary information is missing or contradictory, asking for clarification can be the correct outcome.
Success requires every essential check to pass. For a quote comparison, that means checking both the delivered total and the effective unit price. For a proposed record update, it means preserving the full requested record set. We check the work the contractor needs, down to the outputs that could carry an error into an estimate.
The published Text v0.1 results come from 30 constructed tasks, repeated three times for each of 19 models. They establish what happened on those cases. They do not establish performance across customer projects.
Compare correctness, time, and cost
We report successful attempts, repeatability, elapsed time, and inference cost separately. A quick incorrect answer remains a failure. Failed attempts still consume time and may incur cost.
Direct speed comparisons use tasks both models answered successfully. We show how many tasks qualify and the uncertainty around the result. This prevents a model from appearing faster simply because it left difficult work unfinished. Inference cost means the charge for running the model, not a contractor's software bill or the total cost of preparing an estimate.
Models receive the same inputs in fresh contexts. We retain their settings, outputs, and grading versions so a change can be traced. The active comparison includes models available through Bidlo; when an older model is retired, it leaves new comparison views. Historical results remain dated and intact.
Expand to finished work products
The next step is to test complete deliverables: finding requirements in original plans and addenda, extracting schedules, reconciling quantities and quotes, and producing usable Excel workbooks with correct formulas and complete records.
We reviewed historical requests for contractor work in Bidlo. They included finding pay items across plan sheets, rebuilding quantity tables in Excel, reconciling station ranges, and correcting records while preserving unrelated rows. We are turning these anonymized patterns into synthetic fixtures with file-level checks. This workflow suite is in development and is separate from the published text scores.
Those tests need checks at the file level as well as the answer level. A workbook must open, preserve units and assumptions, and calculate correctly. Plan review must identify the governing revision and point to its evidence. Domain review will help establish whether the checks match the decisions estimators make.
We will version new task sets, rerun eligible comparison models, and document grading changes. Published methods and examples will explain the test while private inputs and expected answers remain protected.