Ideas

Grok 4.7 on Bidlo Bench

Matt Wolfe3 min de lectura

Grok 4.7 passed 86 of 90 attempts on Bidlo Bench's 30 heavy civil estimating tasks, compared with 89 for Grok 4.20 Reasoning. The tasks cover supplied specification changes, bid-item extraction, quantities, delivered prices, and complete record proposals. Its median paired completion time was 0.87× the older model's, with fewer usable, correct results.

Compare the models

Text v0.1 · 30 estimating tasks · 3 attempts each · September 26, 2026

80%85%90%95%100%$0.0000$0.003$0.006$0.009$0.012GPT-6 Astra: 100.0% passed; $0.0053 cost / attemptGPT-6 Astra (Fast): 100.0% passed; $0.0106 cost / attemptGPT-6 Sol: 100.0% passed; $0.0015 cost / attemptGPT-6 Luna: 92.2% passed; $0.000076 cost / attemptGPT 5.6 Terra: 95.6% passed; $0.0013 cost / attemptGPT 5.6 Luna: 96.7% passed; $0.000168 cost / attemptGPT 5.6 Sol: 98.9% passed; $0.0025 cost / attemptClaude Sonnet 5: 96.7% passed; $0.0027 cost / attemptClaude Fable 5.1: 100.0% passed; $0.0088 cost / attemptClaude Opus 5.5: 100.0% passed; $0.0039 cost / attemptClaude Opus 5: 98.9% passed; $0.0065 cost / attemptClaude Haiku 4.5: 98.9% passed; $0.0037 cost / attemptGrok Build 0.1: 96.7% passed; $0.0032 cost / attemptGrok 4.20 Reasoning: 98.9% passed; $0.0033 cost / attemptGemini 3.5 Flash Lite: 98.9% passed; $0.0018 cost / attemptGemini 3.1 Pro Preview: 97.8% passed; $0.0068 cost / attemptTask success · 80–100% scaleAverage inference cost per attempt (USD) · lower is better

Select a point or model row for its measured result and uncertainty.

Each point is one tested model configuration. The line joins observed best trade-offs; it is not a curve of untested reasoning settings.

Rankings by observed contractor-task success
1Adaptive thinking; requested effort: medium100.0%90/90$0.008890/90 billed47590/90 reported4.46 sn=90 passed
1Adaptive thinking; requested effort: medium100.0%90/90$0.003990/90 billed49290/90 reported3.66 sn=90 passed
1Reasoning effort: medium100.0%90/90$0.005390/90 billed30890/90 reported2.22 sn=90 passed
1Reasoning effort: medium100.0%90/90$0.010690/90 billed30890/90 reported1.71 sn=90 passed
1Reasoning effort: high100.0%90/90$0.001590/90 billed34890/90 reported2.25 sn=90 passed
6Thinking budget: 8,000 tokens98.9%89/90$0.003790/90 billed98290/90 reported4.83 sn=89 passed
6Adaptive thinking; requested effort: medium98.9%89/90$0.006590/90 billed55890/90 reported3.32 sn=89 passed
6Requested thinking level: medium98.9%89/90$0.001890/90 billed93590/90 reported4.95 sn=89 passed
6Reasoning effort: high98.9%89/90$0.002590/90 billed32690/90 reported2.16 sn=89 passed
6Provider default; no explicit reasoning setting sent98.9%89/90$0.003390/90 billed1,63690/90 reported7.03 sn=89 passed
11Requested thinking level: medium97.8%88/90$0.006890/90 billed78490/90 reported7.15 sn=88 passed
11Requested thinking level: medium97.8%88/90Unknown88/90 billedUnknown88/90 reported7.27 sn=88 passed
13Adaptive thinking; requested effort: medium96.7%87/90$0.002790/90 billed56690/90 reported3.83 sn=87 passed
13Reasoning effort: medium96.7%87/90$0.00016890/90 billed35090/90 reported2.45 sn=87 passed
13Provider default; no explicit reasoning setting sent96.7%87/90$0.003290/90 billed1,93290/90 reported8.39 sn=87 passed
16Requested reasoning effort: high95.6%86/90Unknown89/90 billedUnknown89/90 reported2.17 sn=86 passed
16Reasoning effort: medium95.6%86/90$0.001390/90 billed32290/90 reported2.04 sn=86 passed
16Provider default; no explicit reasoning setting sent95.6%86/90Unknown88/90 billedUnknown88/90 reported5.92 sn=86 passed
19Reasoning effort: medium92.2%83/90$0.00007690/90 billed35490/90 reported2.02 sn=83 passed

Equal scores share a rank. Small differences on 30 cases do not establish broad model superiority. Costs include failed attempts; incomplete billing or token totals remain unknown. Time is conditional on success. These scores cover supplied-text estimating tasks, not general intelligence or completed plan-to-workbook jobs.

Four failures, two causes

Two Grok 4.7 answers calculated a quote total correctly but omitted a fixed charge from the effective unit price. Both outputs were required. Accepting the correct total alone would have missed an inconsistency an estimator could carry into a comparison.

Two other attempts failed with service errors (Gateway 500) before delivering usable answers. They count against successful delivery, but do not show that Grok misunderstood the construction task. The older Grok model's single failure was an answer that the fixed parser could not read, although its visible arithmetic was correct.

Grok 4.7 passed all three repeats on 27 of 30 tasks; Grok 4.20 Reasoning did so on 29. Compliance with the requested machine-readable format was 82/90 and 79/90 respectively. The newer model followed the requested answer format more often while passing fewer correctness checks. The observed success-rate difference was −3.3 percentage points; its 95% interval, −10.0 to +1.1 points, includes no difference.

Completion time and cost

Both models succeeded at least once on every task, so the paired timing comparison covers all 30. The median candidate-to-baseline ratio was 0.87×, with a 95% interval of 0.81–0.97×. This compares successful repeats of the same tasks.

Billing was returned for 88 of Grok 4.7's 90 attempts, totaling $0.31324. The two errors had unknown cost, so its total and cost per successful result remain unknown. Grok 4.20 Reasoning had complete billing coverage and cost $0.00330 per successful result, including its failed attempt.

This run gives no basis for replacing Grok 4.20 Reasoning on correctness alone. Grok 4.7's shorter observed successful times came with more failed deliveries. Grok 4.20 Reasoning remains available in Bidlo; retired Grok 4.6 is excluded from the charts.

Method

Text v0.1 contains 30 synthetic tasks, repeated three times across 19 models: 1,710 attempts. Models received the same supplied text in fresh contexts, using recorded model settings. No original plan files were parsed, tools used, or Excel workbooks generated. The cases are not a representative sample of customer projects.

Grading v1.3 checks required outcomes and critical errors. After inspecting outputs, we revised grading to accept narrowly defined equivalent representations, applied uniformly to every model. Compliance with the requested answer format remains separate. Task definitions, arithmetic, graders, and results received independent AI review.

Elapsed time spans client request through response parsing. Paired ratios use task-level median successful times; 95% intervals use 2,000 source-group bootstrap resamples. The 180-second execution safeguard bounded requests. Errors retain observed time; timing creates no quality tiers.

Gateway inference cost is the charge to run the model, in USD. It includes unsuccessful attempts with known billing and is not a Bidlo customer charge. See aggregate results and settings, the 19-model comparison table (CSV), and the benchmark introduction.

Archivado en: Ideas

Autor:
Matt Wolfe

Explora el mercado de licitaciones con Licitar.

Elige lo que recibes.

Elige lo que te llega.

  • Producto
  • Investigación
  • Historias