Industrial ops
AI agents operating industrial equipment need a spec nobody wrote
AI agents operating industrial equipment fail on the machine spec, not the protocol. Anthropic's Model Hardware Standard shows what has to exist first.
Anthropic opened a research preview of the Model Hardware Standard on Wednesday, and the coverage settled on the obvious headline: AI can drive lab robots now. The more useful read is buried in the technical description, and it changes what AI agents operating industrial equipment actually require from the operator. The standard’s real deliverable is not a driver. It’s a document — a written statement of what each machine measures, what may be adjusted on it, and what it must never be asked to do. Most plants do not have that document. That is the project.
TL;DR: Anthropic’s Model Hardware Standard, opened as a research preview on 27 August 2026, cuts hardware integration from weeks to hours by standardising two primitives — read and write — and forcing device characteristics and safety limits into a natural-language reference file the driver generates. The speedup is real: Carnegie Mellon reported connecting equipment in about eight hours against the several weeks a vendor-built setup typically takes. But the weeks were never protocol work. They were the cost of reconstructing an operating envelope that lived in a manual, a vendor support thread and one engineer’s head. Anthropic’s standard doesn’t remove that work. It moves it to the front and makes it a file. If you run asset-heavy operations, the envelope is the deliverable, and it is worth writing whether or not you ever adopt the standard.
What the standard actually ships
Strip the demo away and MHS is three plain things.
A driver built on two primitives, read and write — “get temperature”, “set temperature” — so any device that speaks them becomes discoverable in a standard format. No custom translator per instrument.
A metadata layer. Anthropic says the driver “contains tags that let the user write this information directly in natural language” and “then automatically produces a reference file with information about a device’s general characteristics”: weight, measurable quantities, adjustable parameters, safety limits.
And enforcement at the device. One of the pilot researchers describes not needing to worry about an agent accidentally using excess laser power, “because MHS enforces device-level safety limits.” That last one is the design decision that matters. The limit sits under the model, not in the prompt.
The standard is model-agnostic and reachable through existing protocols like MCP. Anthropic says it will open-source it after the preview, once it has built safety evaluations around it.
The integration was never the slow part
Anthropic’s framing of the problem is that it “typically takes a lab or manufacturing facility weeks, if not months, to set up and integrate their hardware.” The pilot numbers against that baseline:
| Organisation | What was connected | Reported result | What it actually measures |
|---|---|---|---|
| Carnegie Mellon | Dose-response experiment setup | About eight hours, “versus the several weeks a vendor-built setup typically takes” | Time to a working integration |
| University of Washington | Six instruments | ”Under a week, including the time I spent writing drivers for them” | Time including driver authoring |
| QuEra Computing | Laser lock recovery on a quantum computer | From 150 seconds and 58% success to about six seconds; across 700 trials it recovered the correct lock 695 times | Performance of one narrow, instrumented routine |
| Tetsuwan Scientific | qPCR liquid handling, 9,143 dispenses across 300 transfer types | Predicted precision 12% more accurately than manufacturer specifications on 31 of 45 runs | Model calibration against a vendor datasheet |
| Genentech | BCA assay, real-time error handling | Water flow around 140 µL/s at 0.016 RMSE; BSA at 10 µL/s, 0.181 RMSE | Control accuracy on two liquids |
Read the first two rows as an operator, not as a scientist. Eight hours instead of several weeks is a factor of about twenty. No protocol change buys a twentyfold improvement in wiring two systems together. What it buys is the elimination of discovery — the part where somebody works out, by trial and by email, what the machine can be told and what happens if you tell it the wrong thing.
I’ve argued the software version of this for a year: the integration is slow because the process was never documented, not because the API is hard. The physical version is the same argument with worse consequences. A wrong write to a CRM field is a bad record. A wrong write to a servo is a bent part, or a person.
Your plant has the same gap, minus the driver
Deloitte’s 2026 Manufacturing Industry Outlook, published November 2025, reports that 22% of manufacturers plan to use physical AI within two years, up from 9% today. The same outlook names the top concern of more than a third of the 600 executives surveyed as equipping workers with the skills and knowledge to get value from smart manufacturing — and suggests agentic AI could be used to capture workers’ tacit knowledge and generate standard operating procedures.
That last suggestion has the dependency backwards, and it’s worth being blunt about it. An agent cannot capture the tacit knowledge of what a machine must not do if the machine has no interface it can read and nobody has written the limits anywhere it can reach. The knowledge has to be extracted by a person, once, and written into something durable. The agent is downstream of that.
Here’s what’s usually true when you go and look. The CMMS has an asset ID, a location and a PM schedule. The historian has tags, many of them named by whoever commissioned the line in 2011. The safety limits exist in three places that disagree: the vendor manual, the PLC, and the informal rule the second-shift lead applies because the manual’s number causes nuisance trips. Nobody is wrong. Nobody owns the reconciliation.
Anyone who has written Modbus code against a vendor’s register map knows the specific shape of this. The map is correct. The scaling factor isn’t in it. The only person who knows the raw value is tenths is the one who commissioned the panel. That is not a data problem you fix with a better model. It’s a page that was never written.
This is one layer in from the map I wrote about in June, when Accenture spent about $4.2 billion on OT asset visibility. Visibility tells you the machine exists. The envelope tells you what it may do. You need both, in that order, and almost every industrial AI proposal I read assumes both already exist.
Where the AI actually earned it
The QuEra result is the honest counterweight, so I’ll give it its due. Laser lock recovery went from 150 seconds and 58% success to roughly six seconds, and across 700 trials the system recovered the correct lock 695 times. That is a serious improvement on a real problem.
Now notice what it is. A single, narrowly scoped, heavily instrumented routine, with an unambiguous success condition, running inside a hard device limit, in an environment where the operating envelope had already been written into the driver. Everything expensive happened before the model touched it.
That’s the pattern, and it’s the same one behind the scheduling engine no off-the-shelf product could build: the win comes from a well-specified subproblem with clean constraints, not from pointing something clever at an ambiguous one.
What a written spec still can’t hold
Anthropic includes a limitation most of the coverage skipped, and it’s the most instructive line in the release. Working with protein samples, the agent produced errors caused by foaming. Anthropic’s account: “Because Claude did not yet understand the underlying physics of the failure, we had to guide it towards parameters that handled the liquid more gently.”
Foaming isn’t in a datasheet. It’s in a technician’s hands. It’s the reason a good operator aspirates slowly on that one buffer and couldn’t tell you why in writing without being asked twice.
So the reference file is necessary and it is not sufficient, and the honest version of the plan says so. You write down everything that can be written down. You enforce the hard limits below the model. Then you accept that the residue — the foaming class of knowledge — surfaces only in supervised running, and has to be written back into the file each time it does. A spec that never gets amended is a snapshot of what one person knew on the day somebody asked them.
The other limitation is more immediately practical. Anthropic states that MHS “doesn’t yet work with hardware that lacks a programming interface.” Walk any older facility and total up the assets that qualify. For those, the standard isn’t the constraint. Serial-to-nowhere is.
Before anyone quotes you for an agent on the floor
Four things, per asset, before a vendor conversation is worth having.
What it measures, at what resolution and how often — and where that stream lands today. What may be changed on it, by whom, and within what range. The hard limits, and specifically whether they are enforced by the device or the controller rather than by an instruction to a model. And the named condition under which it stops, plus the one person who can invoke it without calling a meeting.
That list is unglamorous and it is most of the project. It’s also the artifact that survives your vendor, your model and your platform, which is the whole argument for writing it yourself rather than buying it inside somebody’s product. If you want help getting an existing operation into that shape, that’s the work I do, and the data-contract layer underneath it is usually where it starts.
The standard makes the machine legible to the agent. Someone still has to make it legible to the company first.
FAQ
- What is Anthropic's Model Hardware Standard?
- MHS is a shared specification, opened as a research preview on 27 August 2026, that lets AI agents discover and operate physical devices such as microscopes, liquid handlers and robotic arms. It works through a driver built on two standardised primitives — read commands like get temperature and write commands like set temperature — so a device becomes discoverable in a standard format without a custom translator program for every model. Anthropic says it is model-agnostic and reachable by any agent harness through standard protocols such as the Model Context Protocol, and that it plans to open-source the standard after the preview. Named participants include Genentech, Carnegie Mellon University, the University of Washington Baker and Pinglay labs, HHMI Janelia Research Campus, QuEra Computing and Tetsuwan Scientific, with supporters including AWS, Danaher, Doosan Robotics, QIAGEN, Tecan, Universal Robots and Raspberry Pi.
- Why does connecting AI agents to industrial equipment take weeks?
- Not because the wire protocol is difficult. Anthropic's own framing is that it typically takes a lab or manufacturing facility weeks, if not months, to set up and integrate their hardware. What consumes those weeks is reconstructing information that was never written down in one place: what the device measures, which parameters can be changed, and what its safety limits are. That knowledge is spread across a vendor PDF, a controls engineer's memory and a spreadsheet somebody maintains privately. The integration work is really documentation work, done under a deadline by whoever is available.
- What does an MHS driver actually store about a machine?
- Anthropic describes tags that let the user write the information directly in natural language, and says the driver then automatically produces a reference file with information about a device's general characteristics — weight, the quantities it can measure, the parameters that can be adjusted, and its safety limits. That reference file is the part worth copying even if you never adopt the standard. It is a written, machine-readable statement of what one piece of equipment is and what it must not be asked to do.
- Can AI agents run equipment that has no programming interface?
- No, and Anthropic states the limit plainly: MHS does not yet work with hardware that lacks a programming interface. In a real plant, building or depot, that exclusion covers a large share of the installed base — older drives, standalone controllers, anything whose only interface is a front panel. For those assets the first question is not which agent to buy. It is whether the asset can be read at all, and who owns the tag list once it can.
- What should a manufacturer or facilities operator do before automating equipment with AI?
- Write the operating envelope down, one asset at a time, and give it an owner. For each machine: what it measures, at what resolution and how often; which parameters may be changed and by whom; the hard limits that must be enforced below the automation layer; and the named conditions under which it stops. Then check that the limits are enforced by the device or its controller rather than by the instruction you give the model, because an instruction is a request and a device limit is a wall. This is the same prerequisite as an OT asset inventory, one layer further in: the inventory says the machine exists, the envelope says what it may do.