An independent magazine about AI at work in construction. Every story starts on a real job site, with the people who used the technology and the results they will answer for.
Fable 5.1 Nearly Doubles Anthropic's Workflow Score, With a Safeguard Caveat | ConstructionMagazine.ai
Fable 5.1 Nearly Doubles Anthropic's Workflow Score, With a Safeguard Caveat
Anthropic's September 2026 release scores 31.4 percent on AutomationBench against 17.1 for Fable 5. Part of the gap is a safeguard. Cache reads cost 75 percent less. Zero data retention depends on where you buy it.
Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 in September 2026. On AutomationBench, the benchmark Anthropic lists under the label business workflows, Fable 5.1 scores 31.4 percent. Fable 5 scored 17.1 percent. Opus 5 scored 26.9 percent and GPT-5.6 Sol scored 19.6 percent.
The footnote matters. Anthropic ran the tests with production safeguards on. On tasks where a safeguard intervened, Fable 5 scored zero on AutomationBench. Anthropic writes that this likely reduces the Fable 5 number. The page does not say how many tasks that covered, so the size of the model's own gain is not published. VentureBeat's report on the release put the general caveat plainly: the figures are vendor-reported results, not independent proof.
How the benchmark's task types land on a contractor's desk. AI-assisted illustration.
What Anthropic published about the benchmark
The announcement page gives the label business workflows and the four scores. It does not describe the tasks. The benchmark itself does. AutomationBench was published in April 2026 by Daniel Shepard and Robin Salimans, drawing on workflow patterns from Zapier's automation platform, and Zapier maintains the public repository. Each task starts a simulated business environment across 47 imitation software tools, a CRM, an inbox, a calendar, a helpdesk, a ledger, and asks the agent to leave every system in the correct state. There are 600 scored tasks, 100 each in sales, marketing, operations, support, finance, and HR. A task passes only if every assertion about the final state holds. Partial credit is recorded but does not count toward the pass rate.
Two of those design choices explain why the scores look low. The first is the all-or-nothing pass. An agent that updates the CRM, sends the email, and books the meeting but writes the wrong date into the calendar fails the whole task, which is also how a project engineer's manager would grade it. The second is the private task set. The public repository scores Claude Opus 5 at 50.3 percent and Fable 5 at 46.2 percent at maximum effort, far above the 26.9 and 17.1 percent Anthropic cites. The repository says the official leaderboard runs on a separate, harder, held-out set of tasks per domain. Anthropic's page does not say which set its figures come from, and the two numbers are not comparable until it does.
AutomationBench pass rate, percent, as reported by Anthropic
Anthropic's early-access partners describe the work they tested. Crosby reports its contract redlining benchmark moved from 47.9 to 57.0, with most of the gain on first-turn quality and smaller edits to the documents. Shopify reports long unattended runs where the model keeps its own records and picks up where it left off. Hebbia reports the best fact recall over financial documents of any model it has tested. All three statements were published by Anthropic with the release.
The desk tasks that look like that work
The tasks this magazine has reported from estimating desks have the same shape: many documents, one register, and a person who signs. A bid requirements register pulls dates, bonds, alternates, and insurance limits from the invitation, the specification, and the addenda, and marks a conflict between two limits for a human decision. A drawing revision compare holds two sets still and lists what moved, with a sheet reference on each finding. A subcontractor scope review marks each requirement present, missing, or ambiguous, and does not turn silence into an exclusion.
Each of those tasks fails in the same way when a model is weak at long work: it drops a document, invents a requirement, or resolves a conflict on its own. Contract redlining and unattended multi-step runs are the closest published matches to that failure mode. A benchmark score is not a field result. No contractor has reported a Fable 5.1 job to this desk yet.
What a contractor's version of the benchmark would contain
The six AutomationBench domains have construction equivalents, and each of them has a published cost of doing the work by hand. Submittal review is the largest. Vendors that sell review software put a mid-complexity submittal, a pump or a light fixture package, at two to four hours of an engineer's time and a chiller or switchgear package at four to eight, with project engineers spending twenty hours a week on review in a heavy period. Those are vendor figures with an obvious interest behind them, and they match what estimating desks tell this magazine. A rejected submittal restarts the cycle and adds two to four weeks to the item's schedule.
Lien waivers are the finance domain. A controller running fifteen active projects can have more than 200 waivers outstanding at a month end, each one a document from a subcontractor that has to match a payment before the payment moves. Certified payroll is HR: on a federally funded job, every contractor and subcontractor files a weekly report on Form WH-347 for every worker, including weeks with no work, and one late report from one sub can hold the billing cycle for the whole project. RFIs are support: each one needs a clear question, a drawing reference, and a proposed answer, and a well-run log tracks first-pass approval rates and ageing by trade package.
Every one of those tasks has the AutomationBench shape. Several systems, a policy to follow, data to write into each system correctly, and a pass condition that is binary. An agent that files 199 waivers and misfiles one has not done the job. That is why a pass rate of 31 percent on a hard held-out set is worth reading as good news and as a warning at the same time: the model finishes roughly one such chain in three without a person, and the desk is on the hook for the other two.
Price
Input tokens cost $10 per million and output tokens $50 per million, the same as Fable 5. Cache reads, where the model re-reads context it has already processed, now cost $0.25 per million, a 75 percent cut. Anthropic estimates a typical workload costs about 25 percent less than on Fable 5 and a context-heavy, tool-heavy workload up to about 45 percent less, because cache reads make up most of that bill. A desk that re-reads the same specification book on every check is the second kind of workload.
The arithmetic for a submittal desk is simple to state. A specification book of a few hundred thousand tokens, read once and cached, costs a few dollars to load and a few cents to re-read for each submittal checked against it. The same book read fresh on every check, as a chatbot does, costs ten times more per read. Opus 5, at half the sticker price on every column, is still cheaper per token, and the same Anthropic pricing page shows batch processing at half price again for work that can wait. A desk choosing a model on price should run one week of its own submittals through the cache and read the bill, because the announcement's 25 and 45 percent figures describe Anthropic's typical workloads, not a contractor's.
Cache read price per million tokens, dollars, Anthropic pricing page
Zero data retention
A bid set, a sub's pricing, and an owner's budget go into the prompt. Whether the provider keeps them is a contract question. Anthropic says eligible customers can use Fable 5.1 with zero data retention now, and that Enterprise Frontier Safeguards, which keeps customer data in cloud storage the customer controls, rolls out in phases this fall. Anthropic publishes no eligibility criteria; access follows an existing approved ZDR agreement. A second Anthropic rule cuts across that one. Since 9 June 2026 it retains prompts and outputs from its covered models, the Mythos-class models, for 30 days on every platform, ZDR accounts included. Fable 5.1 is a Mythos-class model. Ask Anthropic which rule applies to your account.
The answer differs by route. On 1 September 2026 the Vercel AI Gateway model list shows Fable 5.1 with zero data retention set to none and no-training set to all. Fable 5 shows the same. Every other Anthropic model on that gateway, from Haiku 4.5 to Opus 5, shows zero data retention set to all. A contractor who reaches Fable 5.1 through that gateway does not have zero data retention today. Ask for the retention terms in writing for the route you use before a bid set goes in.
The Vercel documentation adds a reason for the discrepancy. Its zero-data-retention page, last updated in late July, says the Fable-class model does not support zero data retention on any provider, including Anthropic, Google Vertex, and Amazon Bedrock, because Anthropic requires cumulative visibility across requests to detect some misuse patterns, and that prompts and completions are kept for 30 days and not used for training. Anthropic's own product page says the same thing from the other direction: Fable requires 30-day retention for safety monitoring by default, and only customers eligible for Enterprise Frontier Safeguards get zero retention before that programme arrives. For a contractor the practical reading is that no training is settled and no retention is not. A bid set sent to Fable 5.1 sits with Anthropic for 30 days unless the firm holds an agreement that says otherwise.
What to do with the score
Treat 31.4 percent as a pass rate on chains of small writes, and expect a person to finish two chains in three.
Run one week of the desk's own submittals or waivers through the model with the specification cached, and compare the bill with the current hours.
Get the retention terms for the exact route in writing: Anthropic direct, a cloud provider, or a gateway. They differ.
Keep the writes into the record behind a person's approval until the firm has a month of clean runs on a named log.
Watch for a Fable 5.1 row in the public AutomationBench repository, which is the only number a contractor can reproduce.