Ideas

Grok 4.7 — Bid Bench

Matt Wolfe3 min read

Bid Bench evaluates Grok 4.7 in the context of heavy civil construction and estimating workflows.

The benchmark covers Documents, Takeoffs, Analysis, Knowledge, Apps and Predictions. Bid Bench 1.0 compares delivery quality across thinking settings, with forecasts and stopping behavior scored separately.

Completion limits the efficiency result

Grok 4.7 has one fixed setting in Bidlo. There are no additional app-supported thinking levels to compare.

Bid Bench delivery score20%30%40%50%60%70%80%90%100%0 s10 s20 s30 s40 s50 sGrok 4.7 · Fixed: 68.9% passed; 48.8 sGPT-6 Sol · High: 97.8% passed; 11.7 sGPT-6 Sol · Low: 94.4% passed; 7.4 sClaude Fable 5.1 · High: 93.3% passed; 32.8 sClaude Fable 5.1 · Low: 90.0% passed; 31.5 sGrok 4.20 · High: 75.6% passed; 38.8 sGrok 4.20 · Off: 34.4% passed; 11.1 sGPT-6 SolClaude Fable 5.1Grok 4.20Grok 4.7Median accepted completion time · seconds

All four models shown. Time is measured on accepted work; scores include failures.

Lines connect each model’s best measured trade-offs. All measurements are available in the downloads.

Grok 4.7 at Fixed passes 68.9% of delivery attempts. The comparison settings pass 97.8% for GPT-6 Sol (High), 93.3% for Claude Fable 5.1 (High), 75.6% for Grok 4.20 (High).

Swipe to compare all four models →

Grok 4.7FixedGPT-6 SolHighClaude Fable 5.1HighGrok 4.20High
Cost / 100 accepted———
Delivery score97.8%93.3%75.6%
DocumentsDocument reconciliation88.9%83.3%72.2%
TakeoffsQuantity calculations100.0%100.0%100.0%
AnalysisRecord analysis100.0%83.3%72.2%
KnowledgeConstruction concepts100.0%100.0%66.7%
AppsConstruction functions100.0%100.0%66.7%

High for the three alternatives; Grok 4.7 has one fixed setting. These settings do not represent equal compute budgets. — means incomplete measurement. Full results: CSV · JSON.

It passes 62/90 delivery attempts, or 68.9%. Median accepted completion takes 48.8 seconds. Service and execution failures leave complete cost and token totals unknown; the Cost and Tokens views show measured alternatives. Of the 28 unsuccessful delivery attempts, 22 are service errors, five reach an execution budget and one loses its receipt during a test-host interruption. These are not 28 established reasoning mistakes.


Takeoffs are stronger than Apps

The overall result hides a wide difference between completing schedule calculations and building executable tools.

Grok 4.7 across workflowsFixed0%10%20%30%40%50%60%70%80%90%100%Documents66.7Takeoffs100.0Analysis72.2Knowledge94.4Apps11.1Bid Bench 1.0 · measured September 28, 2026

Documents. Grok passes 12/18 reconciliation attempts. Large document cases contribute to the missing completions.

Takeoffs. It passes all 18 schedule-based quantity checks, including the hardest geometry and material-conversion case. Drawing recognition is outside this edition.

Analysis. It passes 13/18 field-selection and calculation attempts. Historical filters and record revisions are tested alongside simpler queries.

Knowledge. It passes 17/18 construction and estimating scenarios. The remaining attempt is part of the execution failures described above.

Apps. Only 2/18 attempts produce an accepted function. This is the largest practical weakness under the shared five-minute, eight-step budget.


Forecast coverage is incomplete

Both forecasts are scored against recorded outcomes. Zero error is perfect; beating the historical baseline means improving on a simple statistical forecast. The bidder baseline uses past district participation rates, with overall participation as a fallback; the price baseline uses historical median prices for the same item and unit.

Bidder probabilitiesBaseline 6.77ErrorPerfect (0)2.557.5Fixed omitted: 0/12 valid batchesNo complete forecast resultUnit pricesBaseline 35.7%ErrorPerfect (0)20%40%Fixed omitted: 11/12 valid batchesNo complete forecast result

Right of the dotted baseline is better. Bidder error is the Brier score × 100 (0–100); price error is a percentage. The two measures are not directly comparable.

The candidate panels cover 18 of 61 actual bidder/project appearances. Bidder error measures only that restricted panel.

None of Grok’s 12 bidder batches succeeds: the requests end in service errors. Eleven of 12 price batches are valid. On that subset, mean group price error is 37.7%, equal to the historical baseline on the same subset. Comparing it with the full-cohort baseline of 35.7% would be misleading. There is no complete forecast result for either metric.


Harder work exposes the execution limit

The tasks progress from routine work to the hardest cases. The line shows how often Grok completes them correctly at its fixed setting. Grok completes 6/15 attempts at both Frontier and Stress. The easier tiers also have missing completions, so the curve reflects serving reliability and tool execution as well as task complexity.

Tasks completed correctly0%25%50%75%100%123456Fixed · difficulty 1 (Routine): 12/15 completed correctly (80.0%)Fixed · difficulty 2 (Multi-step): 12/15 completed correctly (80.0%)Fixed · difficulty 3 (Complex): 14/15 completed correctly (93.3%)Fixed · difficulty 4 (Expert): 12/15 completed correctly (80.0%)Fixed · difficulty 5 (Frontier): 6/15 completed correctly (40.0%)Fixed · difficulty 6 (Stress): 6/15 completed correctly (40.0%)Fixed 40.0%Task difficulty · 1 easiest, 6 hardest

15 attempts per difficulty level and thinking setting.

All 12 deliberately blocked tasks are identified correctly, with no false stops on solvable work. Median time to a correct stop is 9.8 seconds. That stopping result does not offset the low rate of completed app-building tasks.

ThinkingCorrect stopsFalse stopsStop time
Fixed12/120/909.8 s

False stops abandon solvable delivery tasks. Time is the median for correct stops, excluding timeouts.


Method

Each setting receives 30 delivery cases, repeated three times with frozen inputs, tools and acceptance checks. Cost includes failed attempts; tokens are input plus output per attempt; latency is the median for accepted results. Unknown complete cost or token measurements remain blank. These are controlled text, schedule and function tests, not scanned-plan takeoffs or the full Bidlo app. Method, uncertainty and source corrections · JSON · CSV.


Choosing a setting

Grok is strongest on the supplied takeoff schedules and knowledge questions. Its application-building completion and service reliability limit the broader result. Under this harness, Sol Low delivers more accepted work in less time. Grok needs a more reliable complete run before a cost-efficiency recommendation or forecasting claim is justified.

Filed under: Ideas

Author:
Matt Wolfe

Explore the letting market with Bid.

Choose what you get.

Pick what gets sent.

  • Product
  • Research
  • Stories