An independent magazine about AI at work in construction. Every story starts on a real job site, with the people who used the technology and the results they will answer for.
Blueprint-Bench 2 Scores Floor Plans Rebuilt From Photographs, Not Drawing Sets | ConstructionMagazine.ai
Blueprint-Bench 2 Scores Floor Plans Rebuilt From Photographs, Not Drawing Sets
AI models are getting better at rebuilding apartment plans from interior photos. That score does not prove a model can read a drawing set. Contractors need tests built around revisions and sheet coordination.
Give an AI agent about 20 interior photographs of an apartment and ask it to draw the floor plan. The photos show pieces of the place: a doorway from one angle, a washer and dryer from another, perhaps the same hall seen on opposite sides of a room. The agent has to decide which views belong together and which rooms connect.
That is the assignment in Andon Labs' Blueprint-Bench 2, released in May 2026. It is a demanding spatial-reasoning test with an inviting name for anyone who works in construction. The name can also send the conversation in the wrong direction. No construction blueprint is the input. The model does not open a drawing set, follow a door tag to a schedule, compare two revisions, or reconcile a plan with an elevation.
Blueprint-Bench 2 tests a narrower capability: reconstructing a room-and-door graph from photographs, then turning that inferred structure into a 2D plan. For construction, it is a proxy for one underlying capability.
The benchmark starts from photographs and ends at a plan. AI-assisted illustration.
The assignment is reconstruction, not document reading
Each agent works through 50 apartments in sequence, according to the Andon Labs benchmark page. The agent receives roughly 20 photos for each apartment and keeps a persistent notepad across the run. That memory lets it record a strategy and adjust the strategy on later units.
The requested plan includes rooms, doors, relative sizes, and the overall layout. Scoring reduces the submitted plan and the ground truth to room-connectivity graphs. Rotations and reflections are allowed through D4 symmetry, so an otherwise matching layout is not punished simply because it faces another direction.
The composite score runs from a normalized random baseline of 0 to a perfect result of 1. Half of the score comes from Jaccard similarity on room-to-room connections. Similarity in the number of doors attached to each room contributes 20 percent, and graph density contributes 10 percent. Room count is another 10 percent; door count and orientation receive 5 percent each.
Those weights describe what the benchmark values. They also explain why the headline number should be called a score, not an accuracy rate. A result of 0.386 does not mean that a model produced a plan that was 38.6 percent accurate. It means the output received 0.386 under this normalized, weighted comparison. The score combines several graph properties, and none by itself answers whether a plan is usable for a construction decision.
Andon Labs reports that all models reach about 90 percent on room count. Connectivity is the stronger discriminator. An agent may identify the right number of rooms and still put the wrong door between them, miss a connection, or create one that does not exist. For a superintendent or estimator, that is a familiar distinction: a complete-looking sheet can still carry a consequential relationship error.
The August leaderboard still leaves a wide human gap
The public leaderboard figures below were checked on the Andon Labs page on 1 September 2026. Claude Fable 5.1, released in September, was not on the board on that date.
Entrant — Blueprint-Bench 2 score
Human baseline, 12-apartment subset — 0.586
Claude Fable 5 — 0.386
GPT-5.5 — 0.362
Gemini 3.5 Flash — 0.336
GPT-5.6 Sol — 0.336
Grok 4.6 — 0.332
The human number came from a 12-apartment subset, while agents process 50 apartments, so it is not a matched head-to-head result. Even with that qualification, the gap is substantial: the leading model trails the human baseline by 0.200 on the benchmark's scale.
Andon Labs also reports that several models sit at or below the normalized random baseline. At the frontier, its fitted line rises about 0.06 points per month, with an R-squared of 0.60. That is a leaderboard trend, not a forecast. Model releases are irregular, benchmark entries can change, and a fitted line does not tell a contractor when a system will become dependable on a drawing workflow.
Individual behaviors are more instructive than a claim that models now "understand space." In one example on the benchmark page, Gemini 3.1 Pro used a washer and dryer visible in two photos to work out which way the camera faced. In another, GPT-5.5 inferred from a doorway seen in one image that a bedroom might be a through-room connecting the living area and the hall. These are useful signs of cross-image reasoning. The page gives examples, not an error analysis, so how often either behavior succeeds is not published.
A room-and-door graph leaves most of a drawing set out
A connectivity graph asks which room touches which room through a door. Construction documents carry far more kinds of relationships. A plan may reference a wall type, a detail, a finish note, a section cut, and a schedule entry. Revision clouds may indicate changed work, but the operational question is what changed and what that change affects. Dimensions, scale, symbols, sheet conventions, and trade overlays all matter.
Blueprint-Bench 2 does not claim to test those tasks. Its input is a photo sequence, not a multi-sheet PDF set. Its output is a reconstructed floor plan that can be converted to a graph, not a cited answer tied to a sheet and detail. Its scoring tolerates rotation and reflection because orientation is a small part of the composite. A contractor reviewing a site logistics plan or an equipment-room layout may have no such tolerance.
This boundary does not make the benchmark irrelevant to construction. A model that cannot maintain room connectivity from multiple views would be a weak candidate for photo-based existing-condition work. A model that performs well may deserve further testing. The benchmark can help screen for one underlying capability; it cannot finish the qualification.
The next comparison is with document-centered evaluations. Three exist with published methodology, and each answers a different question:
AEC-Bench — Nomic AI, April 2026. 196 tasks on real public-sector construction documents: drawings, schedules, specifications, and submittals. An automated verifier scores agents against expert answers. Nomic's own agent tops its own leaderboard, so read the ranking as vendor-authored. Its main finding is that retrieval, not reasoning, is the bottleneck: 77 percent of agent runs fell back to plain text extraction.
DrawingVQA — University of Illinois and LSU, CVPR 2026. 92 expert-written questions on 33 issued-for-construction structural drawing sets. Gemini 2.5 Pro scored 71.7, above the average human respondent at 68.4 and far below the professional tier at 94.9.
AECV-Bench — academic, January 2026. Object counting on 120 floor plans and 192 drawing-grounded questions. Text reading reaches about 95 percent. Counting doors and windows stays near 40 to 55 percent.
Scores from different benchmarks should not be placed in one ranking. A graph-similarity score, a visual-question-answering score, and a document-task score measure different outputs under different rules. A model can improve on one while remaining unreliable on another.
Practitioners run their own tests. Tim Fairley's ContractorOS publishes a 169-question drawing benchmark, updated September 2026: nine real drawing sets, 27 models run raw and 12 model-and-harness pairs, binary scoring, disclosed run-to-run noise, and three questions with no answer to catch guessing. The best raw model was GPT-5.6 Sol at 79.1 percent. The best pairing was Claude Opus 5 in Cowork at 91.8 percent. In a June 2026 video Fairley reported an earlier 44-question test of his own: raw images 86 percent at 104,000 tokens per query, extracted text 98 percent at 66,000 tokens, and a structured database of the drawing set 100 percent at 1,400 tokens. Fairley sells the skills that build that database.
Build the test around the decision your team makes
For construction firms, the useful benchmark is a controlled sample of the work the system would handle. A revision-diff test is a strong place to start. Select prior drawing pairs for which the project team already knows the accepted changes. Ask the model to identify additions, removals, and modified requirements, with a sheet and location for every answer. Grade missed changes separately from false alarms. A system that produces a long list of plausible differences can create more review work even when some findings are right.
The sample should include changes that affect downstream work, not only obvious graphic edits. It could test whether the system connects a changed door tag to the relevant schedule entry or follows a revised room designation to another sheet. The exact cases should come from the contractor's own closed projects, with project-specific information handled under the firm's data rules.
A second test should examine cross-sheet retrieval. Give the model a bounded set and ask questions whose evidence spans sheets. Require it to return the supporting sheet numbers and regions. Then check the cited locations, the answer, and whether the model abstained when the set did not contain enough information. This turns an impressive response into something a reviewer can audit.
Keep geometry in its own lane. If the proposed use involves quantities, clearances, or layout, test those outputs against the firm's accepted tools and review process. Blueprint-Bench 2's relative-size component does not establish dimension extraction, scale handling, or takeoff reliability. A room graph can be structurally similar while the geometry needed for the job is wrong.
The acceptance threshold should follow the consequence of an error. A missed finish-note change and a missed life-safety conflict do not belong in the same bucket. Record performance by task and error type, and retain the model version, settings, input set, and output. Frontier model names change faster than construction standards. A result tied to an undocumented configuration is hard to reproduce and easy to overgeneralize.
Read the leaderboard as an invitation to test
Blueprint-Bench 2 gives construction technology teams a public signal about spatial reconstruction. The reported results suggest that leading systems can extract some non-random structure from a sequence of interior views, yet remain below the supplied human baseline. The score says nothing direct about revision review or cross-sheet coordination.
Before a model touches a live drawing workflow, give it old work with known answers. Make it cite the set. Count the misses that matter to the people who would rely on the output. A public spatial benchmark can choose the candidates. The contractor's own test decides whether any candidate belongs on the job.