Ideas

Grok 4.7 on Bid Bench

Matt Wolfe3 min read

Bid Bench evaluates Grok 4.7 in the context of heavy civil construction and estimating workflows.

The benchmark covers Documents, Takeoffs, Analysis, Knowledge, Apps and Predictions. Bid Bench 1.0 compares delivery quality across thinking settings, with forecasts and stopping behavior scored separately.

Completion limits the efficiency result

Grok 4.7 has one fixed setting in Bidlo. There are no additional app-supported thinking levels to compare.

Delivery pass rate70%80%90%100%$0.2$0.5$1$2$5$10GPT-6 Luna LowGPT-5.6 Sol HighUSD / 100 accepted tasks

No complete cost result for Grok 4.7. Available alternatives are plotted; all settings remain in the table.

Bid Bench 1.0 · 30 delivery tasks × 3 runs per setting
Model · ThinkingPassCost¹Tokens²Latency³
68.9% †——48.8 s
70.0%$0.178,4506.9 s
100.0%$9.8321,90617.6 s
All 75 configurations · 18 models
Sorted by observed pass rate; small differences may not separate models.
Model · ThinkingPassCost¹Tokens²Latency³
100.0%$9.8321,90617.6 s
100.0%$11.6626,41916.8 s
100.0%$11.8411,62110.7 s
100.0%$13.9612,84812.8 s
100.0%$15.1525,25821.9 s
100.0%$15.5126,05323.4 s
100.0%$17.4229,68923.3 s
100.0%$17.8713,92118.9 s
98.9%$14.2823,13122.8 s
98.9%——22.2 s
98.9%——27.7 s
97.8%——11.7 s
96.7%$8.7021,28014.4 s
96.7%——17.1 s
95.6%$6.2416,79713.1 s
95.6% †——31.3 s
95.6%——15.8 s
95.6%——13.5 s
94.4%$2.3312,0347.4 s
94.4%$2.9113,38310.4 s
94.4%$4.9319,70510.9 s
94.4%——32.8 s
94.4%——31.5 s
93.3%——32.8 s
93.3% †——31.2 s
93.3%——30.3 s
92.2%——35.5 s
92.2%——33.5 s
92.2% †——29.7 s
92.2%——19.3 s
92.2%——18.2 s
91.1%——35.0 s
90.0%——31.5 s
90.0%——24.5 s
88.9%——35.6 s
88.9%——11.5 s
87.8%——11.5 s
86.7%$4.3417,7379.4 s
85.6%$3.7715,6698.6 s
85.6%——49.3 s
84.4%$7.8419,5669.1 s
83.3%$3.0817,6757.9 s
83.3% †——23.1 s
81.1%——51.7 s
81.1%——39.1 s
81.1%——13.6 s
78.9%$0.2215,8029.6 s
78.9%$0.2515,84410.0 s
78.9%$0.8427,83416.0 s
78.9%——41.3 s
78.9%——66.6 s
78.9%——27.3 s
76.7% †——42.8 s
75.6% †——38.8 s
74.4%——53.0 s
73.3%$0.4612,35413.9 s
70.0%$0.178,4506.9 s
68.9% †——48.8 s
67.8%——67.3 s
65.6%$5.0113,5144.7 s
65.6%——6.4 s
63.3%——26.7 s
63.3%——19.7 s
57.8%$3.2247,65913.3 s
57.8%——36.8 s
56.7%$0.2620,7118.0 s
56.7%$0.5713,5234.6 s
47.8%——78.1 s
46.7%$4.5247,10513.5 s
46.7%$6.4966,16630.4 s
44.4% †——46.2 s
37.8% †——26.8 s
34.4% †——11.1 s
32.2% †——29.3 s
31.1% †——25.1 s

¹ USD per 100 accepted delivery tasks, including failed attempts. ² Mean input + output tokens per attempt. ³ Median seconds on accepted tasks. Missing telemetry is shown as —. Forecasts and judgment are scored separately.

† Includes attempts interrupted by our test host. These do not establish a reasoning error.

Grok 4.7 · Fixed: 62/90 passed. Case-resampling interval 57.8–78.9%. Incomplete cost telemetry; 48.8 s median on accepted tasks. 1 test-host interruption was counted as unsuccessful.

Results JSON · CSV · Method · September 28, 2026 UTC

It passes 62/90 delivery attempts, or 68.9%. Median accepted completion takes 48.8 seconds. Service and execution failures leave complete cost and token totals unknown; the Cost and Tokens views show measured alternatives. Of the 28 unsuccessful delivery attempts, 22 are service errors, five reach an execution budget and one loses its receipt during a test-host interruption. These are not 28 established reasoning mistakes.

Takeoffs are stronger than Apps

The overall result hides a wide difference between completing schedule calculations and building executable tools.

Grok 4.7 · Pass rate by workflow and thinking level

Documents. Grok passes 12/18 reconciliation attempts. Large document cases contribute to the missing completions.

Takeoffs. It passes all 18 schedule-based quantity checks, including the hardest geometry and material-conversion case. Drawing recognition is outside this edition.

Analysis. It passes 13/18 field-selection and calculation attempts. Historical filters and record revisions are tested alongside simpler queries.

Knowledge. It passes 17/18 construction and estimating scenarios. The remaining attempt is part of the execution failures described above.

Apps. Only 2/18 attempts produce an accepted function. This is the largest practical weakness under the shared five-minute, eight-step budget.

Forecast coverage is incomplete

A forecast must return every required batch before we plot a complete error result.

Grok 4.7 · Forecast error against historical baselines; lower is better

The candidate panels cover 18 of 61 actual bidder/project appearances. Bidder error measures only that restricted panel.

None of Grok’s 12 bidder batches succeeds: the requests end in service errors. Eleven of 12 price batches are valid. On that subset, mean group price error is 37.7%, equal to the historical baseline on the same subset. Comparing it with the full-cohort baseline of 35.7% would be misleading. There is no complete forecast result for either metric.

Harder work exposes the execution limit

Grok completes 6/15 attempts at both Frontier and Stress. The easier tiers also have missing completions, so the curve reflects serving reliability and tool execution as well as task complexity.

Grok 4.7 · Pass rate across six difficulty tiers

All 12 deliberately blocked tasks are identified correctly, with no false stops on solvable work. Median time to a correct stop is 9.8 seconds. That stopping result does not offset the low rate of completed app-building tasks.

ThinkingCorrect stopsFalse stopsStop time
Fixed12/120/909.8 s

False stops abandon solvable delivery tasks. Time is the median for correct stops, excluding timeouts.

Method

Each setting receives 30 delivery cases, repeated three times with frozen inputs, tools and acceptance checks. Cost includes failed attempts; tokens are input plus output per attempt; latency is the median for accepted results. Unknown complete cost or token measurements remain blank. These are controlled text, schedule and function tests, not scanned-plan takeoffs or the full Bidlo app. Method, uncertainty and source corrections · JSON · CSV.

Choosing a setting

Grok is strongest on the supplied takeoff schedules and knowledge questions. Its application-building completion and service reliability limit the broader result. Under this harness, Sol Low delivers more accepted work in less time. Grok needs a more reliable complete run before a cost-efficiency recommendation or forecasting claim is justified.

Filed under: Ideas

Author:
Matt Wolfe

Explore the letting market with Bid.

Choose what you get.

Pick what gets sent.

  • Product
  • Research
  • Stories