# Bid Bench 1.0 — measurement protocol

This edition evaluates models in the context of heavy civil construction and estimating workflows. It measures isolated model-and-tool performance under a shared test harness, not the latency or quality of the full Bidlo application.

## Roster and execution

The published roster covers 18 model entries and 75 selectable model/thinking configurations. OpenAI Fast variants are excluded uniformly at the user's request; standard Astra remains included. Provider settings come from the app's runtime resolver. Settings that the app does not expose are not invented. Grok 4.7 and Grok Build have one fixed setting in this roster. Every successful model step must verify its served identity.

Every reported case runs three times per configuration. Frozen inputs and settings make the procedure reproducible; stochastic model responses and live serving conditions can still change on a rerun. A request receives the same source-reading and isolated JavaScript tools, at most eight model steps, at most 32,768 output tokens per model step, and a five-minute wall-clock deadline. Global concurrency is 12, with model order interleaved. Transient service failures permit up to two recorded retries within the same deadline; other failures stay in the denominator. No cross-model fallback is allowed.

JavaScript runs in QuickJS with 32 MiB memory, a one-second CPU budget, and no network, filesystem or process access. The model can read only sources declared for its case. Source reads and calculation calls do not share automatic data bindings: the model supplies calculation inputs explicitly. Large-source results therefore include the ability to transfer and reconcile data within this tool interface. Model answers, tool traces, routing metadata, usage, timing and metered Gateway cost are retained internally. They are not included in the public aggregate downloads.

## Tasks and acceptance

Delivery includes 30 controlled cases: six each for Documents, Takeoffs, Analysis, Knowledge and Apps. The five categories have equal weight. A case passes only when its entire acceptance contract passes; partial answers do not earn fractional credit. Published workflow percentages use 18 attempts per category and configuration. Difficulty tiers run from Routine through Multi-step, Complex, Expert, Frontier and Stress. They describe authored complexity and have not been independently calibrated with human contractors.

Documents reconcile issued amendments, withdrawals, subsidiary payment and field-level cancellation precedence against structured extracted source text. Takeoffs integrate cross-section schedules, overlapping exclusions and reuse zones, then convert volume states and apply truck limits. These cases do not test scanned-plan recognition or scale extraction from drawings.

Analysis tests letting-date selection, local-date filters, approved revisions, composite record identities and historical cutoffs. Knowledge uses closed-answer construction and estimating scenarios. Apps tests executable JavaScript domain functions for order quantities, quote leveling, interval takeoffs, procurement and crew scheduling. Hidden cases compare outputs to deterministic reference calculations. These are component functions, not complete graphical applications.

Numerical tolerances, required keys, exact choice tokens and behavioral tests are set before inference. Execution or identity failures cannot pass. Optional prose is accepted. When a response contains preliminary status objects and a final answer, grading extracts the last complete JSON object containing status. A uniform parser correction prevents an earlier fenced status stub from hiding that final answer; the original strict delivery pass counts are retained in the public data. The Frontier Analysis prompt leaves its projects field untyped. A county project list is therefore accepted as equivalent to a count only when every member exactly matches the expected county membership, without omissions or duplicates. A uniform normalization handles this ambiguity; parser-only and original strict pass counts remain in the download. No numerical tolerance or execution budget changes. No subjective model judge assigns scores.

Four separate stopping cases contain a missing input, an unavailable outcome or an infeasible constraint without disclosing the correct blocker in the question. Correct stops must identify the actual blocker. On the future submitted-price case, either outcome or chronology identifies the same essential blocker; both are accepted uniformly. This prevents an arbitrary choice between overlapping reason codes from being scored as a reasoning error. False stops are blocked answers to solvable delivery cases. Stop time is the median for correct stops, excluding timeouts.

## Retrospective predictions

Bidder forecasts cover 12 TXDOT projects and 20 candidates per project. The candidate universe is selected from pre-cutoff participation; it is not a complete roster of every possible bidder. These panels contain 18 of 61 actual bidder/project appearances (29.5%). Brier error therefore measures calibration within a restricted candidate panel, not recall of every actual bidder. Each target has multiple possible positive bidders, so probabilities do not need to sum to one. Brier score is the mean squared probability error. The baseline uses prior district participation frequency, falling back to overall frequency when district history is absent.

Price forecasts cover 16 project-item targets in four exact item-code/unit groups. The target is the median submitted contractor unit price. Each group's WAPE is the sum of absolute prediction errors divided by the sum of actual prices; the report averages group WAPE rather than pooling incompatible units. The baseline is the median of supplied prior project-item medians in the same code/unit group.

History spans January 2024 through April 2026. Targets fall between May 1 and July 21, 2026. Engineer-estimate pseudo-bidders, unidentified contractors and nonpositive or unlinked bid totals are excluded. The recorded project date must match an explicit letting date. Target outcomes are held outside model-visible inputs. Public graphs show forecast error only when every required batch returned valid predictions; coverage remains in the downloadable data.

This is a small retrospective reconstruction from Bidlo's historical export, not an untouched prospective test. Effective availability of project descriptors at the original prediction date has not been independently verified. Forecast errors should not be read as validated production prediction accuracy.

## Quality and resources

Delivery pass rate is accepted attempts divided by all 90 scheduled delivery attempts for that configuration. Cost per 100 accepted tasks is 100 times all delivery-attempt cost divided by accepted attempts. It includes metered failed attempts. A terminal response stopped at the output or tool budget still has usable cost when all step costs are present. Costs following errors or retries may be incomplete and are not treated as zero. A configuration with any unknown delivery cost has no complete cost estimate.

Tokens are mean input plus output tokens per delivery attempt. Reasoning tokens already included in output are not added again. Latency is median end-to-end seconds on accepted delivery attempts, including tool work. Conditional latency can favor a model that succeeds only on easier cases, so it must be read beside pass rate. Provider caching, live serving conditions and actual account pricing are included; these are not isolated cold-cache measurements.

Pass-rate intervals use 4,000 deterministic case-cluster bootstrap draws, stratified by category, retaining all three repeats within each sampled case. They describe uncertainty on this authored set, not the population of contractor work. Ties and overlapping intervals should not be interpreted as decisive rankings. Several settings still pass all 90 delivery attempts, including the Stress tier. The resulting 100–100 bootstrap interval reflects no observed failures on this small set; it does not establish perfect reliability on unseen work. This edition cannot separate those settings by delivery accuracy.

## Frozen evidence and source corrections

The main execution batch, a uniformly applied Frontier/stopping extension, a corrected forecast batch, and a clarified Analysis batch have separate immutable input manifests. The Frontier extension was frozen before scored main-run outputs were inspected. After leading configurations still reached the ceiling on completed delivery cases, a sixth Stress tier was authored and frozen before any of its inference. It adds one case in every category for every configuration: document effective-date cutoffs, cubic excavation geometry, historical bidder withdrawals and aliases, combined estimating scenarios, and exact stock-cutting procurement. Its reference answers were checked with independent field selection, exact quadrature, an independent SQL query, and two stock-cutting solvers. This later tier is an exploratory extension, not an untouched confirmatory set. A source audit found the original forecast export contained the engineer's estimate as a pseudo-bidder. Those original forecasts are retained but excluded from publication; every configuration receives the corrected inputs and labels. A stopped response exposed an ambiguity in the original Analysis fixtures: the base estimates lacked an explicit approval status. All four affected cases were rerun for every configuration with that premise clarified and the numerical answers unchanged. Their original attempts remain unreported diagnostics. The Frontier Analysis case already specifies approved records. Four originally prompted stopping cases are retained only as instruction-following controls.

The published edition comprises 9,450 reported attempts: 6,750 delivery, 900 stopping and 1,800 corrected forecast attempts. An additional 3,160 executed attempts are retained internally: 2,400 superseded forecasts, ambiguous Analysis fixtures or prompted controls for the published roster, plus 760 attempts from the removed Fast variant. After the source audit, 1,280 unstarted jobs on the superseded forecast cohort were cancelled; another 30 unstarted Fast requests were cancelled when that variant was removed; cancellation records are distinct from executed model calls. Connectivity and tool preflights are also separate from scores. Original outputs are never overwritten by a retry or a grading correction.

The test-host process stopped during twelve started delivery attempts. Their partial traces were preserved, the absent final receipts were recorded as infrastructure interruptions, and those calls were not resent. They remain unsuccessful attempts in the scheduled denominator; they are not evidence of reasoning errors. Each affected configuration is flagged with an infrastructure_interruptions count in the downloads. Complete cost, tokens and duration for these attempts are unknown. Conclusions about small differences involving these configurations require that limitation.

Validation consists of automated reference checks, frozen-hash verification and primary-agent evidence review. No independent human domain-expert review is claimed. Public JSON and CSV provide measured aggregates; input snapshots, private labels and full execution traces are archived internally for reproduction.
