An independent magazine about AI at work in construction. Every story starts on a real job site, with the people who used the technology and the results they will answer for.
Drawing Benchmarks Agree: Models Read the Sheet but Fail the Count | ConstructionMagazine.ai
Drawing Benchmarks Agree: Models Read the Sheet but Fail the Count
Five tests published since August, by vendors and universities, score frontier models on real construction drawings. Schedules and notes come back near human. Counting and scale measurement do not. The numbers side by side.
Five separate tests of AI models on construction drawings were published or updated between 1 August and 2 September, by a takeoff vendor, an estimating-skills vendor, two university groups, and a document-AI firm. They use different drawing sets, different scoring, and different models. They reach the same result. A frontier model reads a schedule, a note, or a title block almost as well as a person. Ask it how many footings are on the sheet, or how long a wall is when the number is not printed, and it fails at a rate no estimator would accept from a first-year hire.
This article puts the five tests side by side, with the numbers each one published and what each one leaves out. Two of the five are run by companies that sell the thing being tested. That is noted where it applies. The pattern holds across the other three.
ContractorOS: 169 questions, 27 models, and a counting column
The largest public table is the ContractorOS benchmark, updated in September by Tim Fairley's ConstructIQ. An estimator wrote 169 questions against nine real drawing sets spanning architectural, structural, electrical, mechanical, and plumbing sheets. The questions went to 27 models through the raw API with no tools, and then to 12 model-and-harness pairs, meaning the model ran inside a product such as Claude Code, Cursor, Codex, or ChatGPT Work with the drawings in a folder it could open. Scoring covers the 134 questions the drawings can answer. An answer is right or wrong, and a partly right answer to a three-part question is wrong. Three questions cannot be answered from the drawings, and saying so is the correct answer.
The best model alone was GPT-5.6 Sol at 79.1 percent, at $13.54 for a full run. The best pairing was Claude Opus 5 inside the Cowork harness at 91.8 percent, 23.9 points above the same model on the raw API. Claude Fable 5 alone scored 65.7 percent. Grok 4.6 alone scored 49.3 percent and 79.8 percent inside Cursor, the largest lift in the table. The page warns that a repeat run moves by 1.1 to 5.4 points on its own, so anything closer than that is a tie.
GPT-5.6 Sol on the raw API, percent correct by question type (ContractorOS, September 2026, 134 answerable questions)
The column that matters for takeoff is the split by question type. GPT-5.6 Sol, the best raw model, got every schedule question and every scope question right and 93 percent of the plain sheet-reading questions. On counting it scored 64.7 percent across 34 questions. On scale and measurement, where the number is not printed and the model has to work it off the drawing, it scored 36.4 percent across 11. The Claude models were worse on those two columns: Fable 5 at 32.4 percent on counting and 9.1 percent on scale, Opus 5 at 35.3 percent on counting with a fifth of its answers declined. The page calls counting the core takeoff skill and the one most people want to hand over first.
Two caveats. ConstructIQ sells drawing-reading skills and a paid community, so the vendor is grading the market it sells into. And the September table supersedes a nine-model video Fairley posted on 15 August with slightly different figures, 81.3 percent for GPT-5.6 Sol and 44.7 percent for Grok. The drawing sets are not public, so nobody outside the company can rerun it. The harness results are one session per pairing, run once, which the page itself says is inside the noise band. Read the raw-API column as the durable result and the harness column as a direction.
Handoff: a takeoff agent scored against working estimators
On 15 August Handoff AI posted a paper on arXiv describing TakeoffBench-v1: ten real residential blueprint sets, permission obtained and personal data stripped, paired with expert takeoffs of 2,009 line items, of which 1,348 primary materials are scored. Nine trades are covered, from concrete and framing to plumbing and windows. Scoring is per trade, on coverage and on quantity precision within 25 percent of the gold figure, by an LLM judge run ten times per prediction set.
The result that travels is the comparison against people. Independent professional estimators scored against the same reconciled answer key reached a composite of 77.6, with 65.5 percent coverage and 87.9 percent precision. Handoff's own agent, H1, reached 81.6, with 86.1 percent coverage and 78.8 percent precision. Seven frontier and open-weight models working from the raw PDF spanned 35 to 61. Claude Fable 5 was the best of them at 61.4, GPT-5.6 at 58.6. The agent led five of nine trades, the estimators led siding decisively and electrical and windows by small margins, and no frontier model led any trade.
TakeoffBench-v1 composite score, residential blueprint takeoff (Handoff AI paper, 15 August 2026)
Read the two halves of the human score separately. The estimators found fewer items than the agent but were far more precise on the ones they found. The agent found more and was wrong on quantity more often. A contractor deciding which failure is cheaper should notice that a missed item is a change order and a wrong quantity is a bid. The paper is written by the vendor about its own product, the answer key is available only on request, and the sets are residential. Handoff's own X post on 24 July gave H1 a score of 85.2 percent on an earlier version of the test. The paper says 81.6. Both are Handoff's numbers.
Reading the schedule passes. Counting the footings is where every August test lost points. AI-assisted illustration.
DrawingVQA: professionals at 95, models at 42 on the takeoff column
Yoonhwa Jung, Junryu Fu, and Mani Golparvar-Fard posted DrawingVQA on arXiv on 16 July, two weeks before the window this article covers, and it is included because it is the only academic test with an explicit takeoff column. The set is 33 issued-for-construction structural drawings from six real campus building projects, with 92 expert-written questions at three depths: what is on the sheet, what it means in context, and what a domain expert would conclude.
Experienced professionals scored 94.9 overall and 96.6 on quantity takeoff. The average human scored 68.4. The best model, Gemini 2.5 Pro, scored 71.7 overall, above the average person, and 41.7 on quantity takeoff, with several models at zero on that column. The models tested are a generation old, which cuts both ways: newer models would score higher, and the takeoff gap in this test is the same gap ContractorOS measured on this month's models.
CrossProjection: right answer, wrong place
The fourth test explains the mechanism. CrossProjection, posted on arXiv on 1 August by Kaho Li and colleagues, asks whether a vision model can keep track of the same building component across plan, section, and elevation and then point to where it is. Across 23 real drawing sets and 1,954 conditions, GPT-5.5 answered the categorical questions correctly 82.4 percent of the time. Three architecture-trained people scored 87.3 to 93.3. When the same model had to output the geometry, a point, a region, or a line endpoint, the point and region hit rates dropped to between 54 and 76 percent and the line-endpoint hit rate to 22 percent.
That is the counting problem seen from the inside. A model can say the right thing about a footing and still not know where the footing is on the sheet. Counting and measuring both require the second skill. The paper's conclusion is narrow and worth carrying: success on a multiple-choice or marked-element question is not evidence that the model can locate anything.
Two vendors, two ground truths
Two more August results come from companies scoring themselves. Kreo published Auto Measure 3.0 on 28 August with outline accuracy up from 72 to 94 percent and area captured from 92.6 to 99.3 percent against a labelled ground truth of noisy plans. It names no models and releases no set. AECFoundry posted on 25 August that its system scored 87 percent on Nomic's AEC-Bench, a 196-task set of drawing and specification tasks, with 96.6 percent on multi-sheet reasoning and 75 percent on submittal review against 23.1 for the best generic configuration. Its founder's line is that on real construction documents the gap between systems lives in the extraction layer, not the agent layer. Both are consistent with the other tests, and both are marketing.
What every August drawing test found, in the order the skills fail
1Read a schedule or a note: near human, 93 to 100 percent for the best model
→
2Follow a reference from sheet to sheet: 77 to 96 percent
→
3Count the elements on a plan: 32 to 65 percent for frontier models alone
→
4Measure from scale when the number is not printed: 9 to 36 percent
→
5Give a takeoff quantity within 25 percent: 35 to 61 on Handoff's composite
What a contractor can do with this
The tests agree on the shape of the tool, so the buying question is what to put around it. The ContractorOS table shows that the same model inside a harness that indexes the set first gains 11 to 30 points, and the largest gains go to the weakest raw models. The Handoff paper shows a purpose-built agent beating both frontier models and, on coverage, the estimators. Neither closes the counting column on its own. An estimator who lets a model read the specifications, pull the schedules, and trace the cross-references, and then counts the footings themselves, is using the part that works and keeping the part that does not.
Three things to ask any vendor showing a drawing score this month. Which questions are counting and measuring, and what is the score on those alone. Is the drawing set public, and if not, who wrote the answer key. And was the run repeated, because the one public table that measured it found that a single run moves by up to five points on its own.