An independent magazine about AI at work in construction. Every story starts on a real job site, with the people who used the technology and the results they will answer for.
Grok 4.7 Leads SpaceXAI's Electrical and Legal Benches at Grok 4.6 Prices
SpaceXAI released Grok 4.7 on 21 September 2026. Its own table puts the model first on EEBench and the Harvey legal bench at $2 and $6 per million tokens, while Fable 5.1 keeps the coding leads.
SpaceXAI released Grok 4.7 on 21 September 2026, and the figure that matters most to construction readers sits in a corner of the company's own comparison table rather than at the top of it. On EEBench, an electrical engineering bench, SpaceXAI reports Grok 4.7 at 64.0 percent, ahead of Fable 5.1 Max at 56.4 percent, Grok 4.6 at 53.0 percent, and GPT-5.6 Sol Max at 39.4 percent. On the Harvey Legal Agent Benchmark the company reports 19.6 percent for Grok 4.7, with the other three models further down. API pricing holds at $2 per million input tokens and $6 per million output tokens, unchanged from Grok 4.6, according to OfficeChai's report of the announcement. For a general contractor, that is a reason to run a cheap trial on electrical submittals and on contract triage, within the limits set out below.
As of 21 September 2026, this article rests on the SpaceXAI launch comparison table, same-day coverage by OfficeChai and Decrypt, and the Artificial Analysis model page for Grok 4.7. Every benchmark score from the table is company-reported and unaudited. OfficeChai adds that CursorBench is built by Cursor and that independent evaluations are still pending. No source reports a contractor running Grok 4.7 on a project, and no source gives a measured construction result. The scores below are bench scores and nothing more.
Stock office desk with laptop and rolled plans. Library reuse after generate_article_image not attempted for this draft.
The table SpaceXAI published
SpaceXAI's comparison sets Grok 4.7 at its xHigh effort setting against Grok 4.6 at High, GPT-5.6 Sol at Max, and Fable 5.1 at Max, per OfficeChai. The research corpus for this article says the table covers seven benches and carries figures for six of them: EEBench, the Harvey Legal Agent Benchmark, CursorBench 4.0, DeepSWE v1.1, AA-Briefcase v1.1, and Terminal-Bench 4.0. The effort settings differ across the columns, which is normal vendor practice and also a reason to treat the gaps as indicative rather than exact. A model run at a higher effort setting spends more tokens per answer, and the table does not report cost per bench.
The electrical result is the clearest lead in the table. Grok 4.7 at 64.0 percent sits 7.6 points above Fable 5.1 Max and 11.0 points above the previous Grok, on SpaceXAI's numbers. The legal result is a larger relative gap on a much lower base. Grok 4.7 at 19.6 percent is ahead of Grok 4.6 at 15.8 percent, Fable 5.1 Max at 6.7 percent, and GPT-5.6 Sol Max at 2.5 percent, again on SpaceXAI's numbers. Both figures below are drawn from that table and carry the same company-reported label.
21 September 2026. Supports the release date, the $2/$6 per million token pricing unchanged from Grok 4.6, the fast variant at twice the price, the four-model effort settings, the per-bench scores as company-reported, and the caveats that CursorBench is built by Cursor and independent evaluations are pending. Article-level URL to be confirmed by the editor; the research corpus carries the publisher and date only.
21 September 2026. Supports the release date, the pricing, the GDPval Elo of 1695 for Grok 4.7 against 1735 for Fable 5.1, and the company pitch that SpaceX engineering data (Starlink telemetry, manufacturing records, failure logs) went into training. The training-data claim is attributed as Decrypt's report of the company's pitch. Article-level URL to be confirmed by the editor.
Read 21 September 2026. Supports the Intelligence Index of 46 for the xhigh setting, the 500k context window, the 21 September 2026 release date, the creator label SpaceXAI, and the listing of AA-Briefcase among index evaluations.
Free for construction companies.
One email a week on AI in use on real projects, from the contractors and people who ran it.
The SpaceXAI table describes EEBench as an electrical engineering bench. Neither OfficeChai nor Decrypt describes the task set, so this article cannot say whether the bench covers code questions, circuit analysis, equipment sizing, or textbook problems. That matters because the electrical document work a general contractor pays for is narrower and messier than any of those. It is checking a switchgear submittal against Division 26 of the spec, or answering an RFI about a feeder that does not match the one-line diagram. A bench score says the model handles some electrical reasoning better than its peers under test conditions. It does not say the model reads a scanned as-built from 2019 without error, and no source claims it does.
Decrypt reports that xAI and SpaceXAI folded SpaceX engineering data into training, including Starlink telemetry, manufacturing records, and failure logs. That is Decrypt's account of the company's pitch, and it should be read that way. Aerospace manufacturing data is not building electrical data. A model trained on failure logs from a rocket factory may reason well about circuits and still know nothing about a local amendment to the electrical code. The practical test for an MEP coordinator is the same as it was before this release: put the model on a set of documents the team has already reviewed, and count where it disagrees with the reviewer.
Legal scores stay low, including for the leader
The Harvey Legal Agent Benchmark result deserves the most careful reading in the table. Grok 4.7 leads at 19.6 percent on SpaceXAI's numbers, but on the bench's own scale that is under one in five. The research corpus for this article does not describe how Harvey scores a task or what a passing answer looks like, so the figure cannot be translated into an error rate on contracts. What can be said is that every model in the table, including the leader, scores low, and that the next model down is the previous Grok rather than either rival flagship. A relative lead of that kind is a reason to test, and no reason to trust.
For a contractor, the legal use cases are contract triage rather than legal advice. Sorting incoming correspondence by whether it carries notice language, or pulling every liquidated damages clause out of a subcontract package, are review tasks with a human decision at the end. A model that scores under 20 percent on a legal-agent bench belongs in the first pass of that work, with a contracts manager or counsel reading every output before anything is sent. The same applies to change orders that reference a clause the project team has not seen and to RFIs that carry claim language. No source reports a contractor doing any of this with Grok 4.7, and this article does not claim one has.
Where Fable 5.1 keeps the lead
The same SpaceXAI table shows why this release is a niche story rather than a general one. On CursorBench 4.0, Fable 5.1 Max scores 51.8 percent to Grok 4.7's 46.3 percent. On Terminal-Bench 4.0, Fable 5.1 Max scores 57.9 percent to Grok 4.7's 38.0 percent, a gap of nearly 20 points. On AA-Briefcase v1.1, Fable 5.1 Max posts 1678 to Grok 4.7's 1657. On DeepSWE v1.1, GPT-5.6 Sol Max leads at 72.7 percent with Grok 4.7 at 71.0 percent, and the Grok figure carries an asterisk in the table marking it as a High Effort run rather than xHigh. All of these are company-reported by SpaceXAI, and OfficeChai notes that CursorBench is built by Cursor, which is also a distributor of the model.
Coding benches where Fable 5.1 leads, percent, SpaceXAI launch comparison table (company-reported, unaudited), 21 September 2026
Two outside numbers point the same way. Decrypt reports a GDPval Elo of 1695 for Grok 4.7 against 1735 for Fable 5.1. Artificial Analysis lists Grok 4.7 at the xhigh setting with an Intelligence Index of 46 and a 500,000-token context window, records the release date as 21 September 2026, and labels the creator as SpaceXAI. Artificial Analysis also lists AA-Briefcase among the evaluations in its index, which is the one place where an independent tracker and the vendor table overlap. The research corpus for this article does not include an independent run of EEBench or Harvey from any tracker, so the two construction-relevant leads rest on the vendor's numbers alone as of publication.
Price is the routing argument
Grok 4.7 costs $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6, per OfficeChai and Decrypt. A fast variant runs at twice that price for twice the output speed, per the same reports. The model is live in Cursor, Grok Build, the Grok API, and cloud platforms. The research corpus for this article does not carry list prices for GPT-5.6 Sol or Fable 5.1, so the cost comparison across vendors is left to the reader's own contracts. Holding the price flat while the electrical and legal scores rise is what makes a trial cheap. A firm that already pays for Grok 4.6 can swap the model string and rerun last month's document set for the same token bill.
That is the construction software angle. A general contractor's software stack already routes work by type, whether or not anyone calls it routing. Estimating, document control, scheduling, and contract administration tools each call whatever model their vendor chose. Where a firm controls the model choice itself, through an internal tool or a platform that exposes model selection, the SpaceXAI table argues for a split: keep the model that leads the coding and terminal benches on automation work, and trial Grok 4.7 on electrical documents and contract language, where it leads the same table. The flow below is one way to structure that trial. It is a suggested process, not a report of any firm's practice.
How a GC could route a trial of Grok 4.7 by task type, based on the SpaceXAI comparison table of 21 September 2026 (suggested process, not a reported practice)
1Incoming task or document
→
2Classify it: electrical or MEP document, contract or RFI language, or coding and terminal automation
→
3Electrical or MEP: run Grok 4.7 and the current model on the same already-reviewed document set
→
4Contract or RFI: run Grok 4.7 as a first-pass sorter with a contracts manager or counsel reading every output
→
5Coding or terminal automation: keep the model that leads CursorBench and Terminal-Bench in the table
→
6Log tokens per document and disagreements with the human reviewer for each model
→
7Decide routing from the logged counts, not from the vendor table
What to ask before putting Grok 4.7 on electrical or contract work
Ask what EEBench tests. If SpaceXAI or an independent tracker publishes the task set, compare it against the documents the electrical team handles every week. A bench built on circuit theory says little about submittal review, and a bench built on code questions says little about reading a panel schedule from a scan.
Ask for the effort setting and the cost per document. The 64.0 percent figure comes from the xHigh setting, per the SpaceXAI table, and a higher effort setting spends more output tokens per answer. The $6 per million output price is only cheap once the token count per document is known, and the vendor table does not report it.
Ask who reads the legal output. On a bench where the best score in the table is 19.6 percent, the model's role is to sort and flag. Decide before the trial which person signs off on each flagged item, and log every case where the model missed a clause the reviewer caught. That log is the only evidence the firm will have if a missed notice deadline is ever traced back to the tool.
Ask whether the documents leave the building. Contract packages and electrical drawings are often covered by confidentiality terms in the prime contract. Confirm where the API sends the documents, what the provider retains, and whether the cloud platform the firm already uses offers the model under its existing data terms before the first upload.
Ask for a control. Run the current model and Grok 4.7 on the same reviewed document set, count the disagreements with the human reviewer for each, and keep the receipts. That count, and not the SpaceXAI table, is the number that should decide the routing. If the count is never produced, the firm has bought a benchmark headline and nothing else.
Grok 4.7 Leads SpaceXAI's Electrical and Legal Benches at Grok 4.6 Prices | ConstructionMagazine.ai