Enterprise AI
The AI agent harness sets the bill. Nobody in your org chose it.
An AI agent harness is everything in an agent except the model. It moved cost 2× to 5× with no change in task success, and it's the layer you actually own.
An AI agent harness is everything in an agent system except the model. The instructions, the tool registry, the context it assembles, the permissions, the sandbox, the memory, the log. Three September releases from three very different places say the same thing about it: a university lab, a model vendor and a signals agency. Now a cloud provider sells it by name.
On 11 September 2026, the Australian Signals Directorate published guidance on agentic AI harnesses, written for executives and CISOs, that treats the harness as the component an organization can most directly govern. On 16 September, UC Berkeley’s Sky Lab published HarnessTax, the first controlled comparison I’ve seen of the same model running inside different coding-agent harnesses. On 17 September, OpenAI announced Astra for Law, which is GPT-6 Astra with a search index, a set of instructions and 26 plugins wrapped around it. The Berkeley result: the harness moves cost by 2× to 5× without moving task success. The OpenAI result: the harness moves task success by 40% relative, with the model held constant. Same lesson three times. What sits around the model is the variable, and in most organizations nobody chose it.
TL;DR: An AI agent harness is every part of an agent system other than the model, and it is the part your organization controls. Australia’s ASD said exactly that in its 11 September 2026 guidance on agentic AI harnesses, adding that many of the decisions that set an agent’s reliability, operating cost and security risk are made in the harness, not in the model. UC Berkeley Sky Lab’s HarnessTax study (16 September 2026) ran seven models through three harnesses — Claude Code, Codex CLI and Pi, on SWE-bench Lite and Terminal-Bench 2.0, and found the harness moved success rates by roughly ±2% to ±5% while moving cost by up to 5×; Claude Fable 5 solved 97.8% of attempts in Claude Code and 96.7% in Pi at $1.33 versus $0.67 per attempt. The gap is set on the first call: Claude Code’s mean initial context was over 10× Pi’s, from 23 declared tools and roughly 77,000 characters of tool schema. OpenAI’s Astra for Law (17 September 2026) shows the other side: the same GPT-6 Astra goes from 38.7% to 54.0% on Vals AI’s Legal Research Bench once the harness carries a legal index, instructions and integrations. The harness is a spend policy, a security policy and an integration policy at once. It ships as a vendor default, and it needs an owner.
The first controlled experiment on the layer everyone skips
The Berkeley setup is small and clean. Twenty-one model–harness pairs: Claude Fable 5, Opus 4.8, Sonnet 4.6, Haiku 4.5, GPT-5.6 Sol, GPT-5.6 Luna and Kimi K3, each run inside Claude Code, Codex CLI and Pi. Thirty randomly sampled tasks from each of SWE-bench Lite and Terminal-Bench 2.0, three attempts per task, a 100-turn cap, each harness on its own high-effort setting, and one direct-API price list dated 1 September 2026 applied to every model regardless of harness. Cost and success averaged per task, then across tasks, with bootstrap confidence intervals.
The headline, in the authors’ words: harness choice “has little effect on task success rate, but can significantly affect the cost.” Fable 5 solved 97.8% of attempts in Claude Code, 96.7% in Codex and 96.7% in Pi. Claude Code cost about twice Pi — $1.33 against $0.67 per attempt. Across the shared models, Claude Code ran about 2.0× Pi and 1.6× Codex on SWE-bench Lite, and 1.5× Pi on Terminal-Bench 2.0, using geometric means of the cost ratios. The harness effect on success stayed inside ±2% on SWE-bench Lite and roughly ±5% on Terminal-Bench 2.0.
The published chart data is more pointed than the prose. For Fable 5 on SWE-bench Lite, Claude Code and Pi produced the identical outcome on 28 of 30 tasks, and Claude Code was the more expensive harness on 29 of 30. The cost difference survives the study’s multiple-comparison correction. The success difference does not. For GPT-5.6 Luna, Claude Code cost 5.1× Pi — costlier on every one of the 30 tasks — with the same result on 28 of them.
Turn counts don’t explain it. Pi and Claude Code averaged 15.4 and 15.3 turns per attempt for Fable 5 on SWE-bench Lite. Same number of round trips, twice the bill, 1.1 points of success.
And the third finding, the one that should bother procurement: across the six Anthropic and OpenAI models and both benchmarks, a harness other than the model vendor’s own produced the highest observed success rate in nine of twelve comparisons. Sonnet 4.6 did better in Codex than in Claude Code (68.9% versus 66.7%). GPT-5.6 Sol on Terminal-Bench 2.0 did better in Pi than in Codex (83.3% versus 78.9%) at about half the cost ($0.42 versus $0.76). The authors are careful about the limits — two public benchmarks the models may have seen in training, thirty tasks each, results may differ on other workloads. Arena co-authored the study and sponsored the API access. Take the ratios as the shape, not the decimal.
The bill is set on the first call
Where does 2× come from if the turn count is the same? The harness decides what the model reads before it reads your task.
Berkeley measured the first main model call in every SWE-bench Lite attempt: 630 attempts per harness, across all seven models.
| Harness (version tested) | Tools declared on call one | Tool-schema size | Instruction size | Mean first-call context (tokens) | Cost vs Pi, SWE-bench Lite |
|---|---|---|---|---|---|
| Pi 0.85.1 | 4 | ~2,900 chars | ~2,500 chars | ~1,970 | 1.0× |
| Codex CLI 0.146.0 | 7.4 (3–9 by model) | ~18,100 chars | ~23,500 chars | ~11,300 | ~1.3× (Fable 5); Claude Code ran ~1.6× Codex |
| Claude Code 2.1.224 | 23 | ~77,000 chars | ~13,500 chars | ~27,000 | ~2.0× (geometric mean); 5.1× for GPT-5.6 Luna |
Means from HarnessTax’s published first-call chart data (SWE-bench Lite, n = 630 per harness). First-call context is provider-reported prompt tokens including cache creation and cache reads. Cost ratios from the study’s harness-effect data; Codex ratios shown where the study states them.
Twenty-three tool definitions, about 77,000 characters of JSON schema, sent on the first call — and, because a harness re-sends its declarations, on every call after it — whether the task touches those tools or not. Pi sends four: read, write, edit, bash. Portkey’s Siddharth Sambharia measured the same shape five months earlier, in April 2026, by routing the three agents through Portkey’s gateway and reading the raw requests: roughly 2,600 input tokens per request for Pi, 15,000 for Codex, 27,000 for Claude Code, and 83,000 tokens versus 8,000 to write a Fibonacci script. Portkey sells that gateway, so weigh it accordingly. It agrees with the academic number to within rounding.
Prompt caching softens this. It doesn’t remove it. Cache reads are still billed, and the harness — not you — decides what sits above the cache breakpoint. A version bump that adds two tools to the schema invalidates the cached prefix for every session. That’s a cost change nobody approved, in the same way a model-routing default is a business rule nobody wrote down.
OpenAI just sold the harness as the product
Read Astra for Law as an engineering document rather than a launch. OpenAI describes it as GPT-6 Astra “with settings, tools, and context tailored for professional legal work”: a legal search index over more than 230 million URLs and CourtListener’s collection of more than 99.9% of published U.S. precedential case law, custom instructions for legal analysis and writing, and 26 partner-built plugins connecting ChatGPT to iManage, Intapp, DeepJudge, Thomson Reuters HighQ, Relativity and Clio, plus nine community plugins and 47 custom skills.
Then the number. On 200 questions from the private validation set of Vals AI’s Legal Research Bench, at the highest reasoning effort for both systems, the configured product passed the overall correctness check on 54.0% of questions, against 38.7% for GPT-6 Astra using web search alone. OpenAI calls it a 40% relative improvement. It’s the same model. The index, the instructions and the plugins are the whole delta. The firms quoted — Sullivan & Cromwell’s agreement analyzer, Ropes & Gray’s deal-diligence system, Cooley’s IPO tool — were built by OpenAI’s forward-deployed engineers around each firm’s own precedents and playbooks, which is the vendor admitting again that the model was never the bottleneck.
Berkeley shows a harness can double the cost without touching the outcome. OpenAI shows a harness can lift the outcome by 15 points when it carries the domain’s tools and rules. Both are the same variable. The buyer’s mistake is evaluating the model and inheriting the harness.
A signals agency wrote the same definition
ASD’s document is the first government guidance I’ve seen that separates the harness from the model and then tells executives to govern the harness. Its definition: the harness is “the components of an agentic AI system other than the LLM itself.” Its key takeaways, in order: organisations control the harness, not the LLM; the harness is where long-term value, governance and investment accumulate; and the harness is the primary focus for security and risk management.
The sentence that matters for this article is the one about cost. ASD writes that many of the decisions that determine an agentic system’s “reliability, operating cost and security risk are made in the harness, not in the LLM.” That is Berkeley’s finding, stated as policy by a security agency that had no benchmark to sell. ASD goes further than I expected on spend. It recommends cost controls on model APIs as a security measure, because agentic sessions “cost substantially more” than a chatbot exchange, and watching consumption can catch misuse, runaway agent loops or what it calls denial-of-wallet attacks. A token bill nobody reads is a monitoring gap.
The document also lists eleven harness components and calls them “the organisation’s configuration surface.” That list is the most useful thing in it, because it’s a checklist of decisions that already exist in your deployment whether or not anyone made them:
| Harness component (ASD, Sep 2026) | What it decides | Where it shows up |
|---|---|---|
| Prompt and policy layer | System instructions, business rules | Fixed context on every call; behavior drift at version bumps |
| Context manager | What the model sees, what’s dropped | Token bill; sensitive data in the window |
| Tool registry | Which tools exist and how they’re described | Schema weight on call one; attack surface |
| Permission system | Which actions need approval | Security; how often a person is interrupted |
| Execution environment | Where actions run and within what bounds | Blast radius |
| Connector layer | Which enterprise systems, RAG stores and MCP tools are attached | Integration scope; data exposure |
| Memory and session store | What persists between sessions | Stale or manipulated content carried forward |
| Audit and observability | What is recorded | Whether you can explain a cost spike or an incident |
| Update and supply chain path | How new functionality arrives | Unreviewed change to all of the above |
| User interface, model interface | What users see and approve; which models are called | Oversight; routing |
Components from ASD’s “Agentic AI harnesses: the layer above the model” (11 September 2026). The last two columns are my mapping, not ASD’s.
Two of ASD’s practical recommendations are worth lifting straight into a runbook. Where context has to shrink, delete stale content rather than summarizing it, because a summary rewrites the record and can introduce new errors. And keep a persistent, organization-level rules file the harness reads every session: standards, prohibited areas, build and test commands. That rules file is the harness configuration written down. It’s the artifact this article has been arguing for, and a signals agency now recommends it as security practice.
Then the market did the obvious thing. DigitalOcean’s Managed Agents, in public preview and covered by InfoQ on 2 October 2026, has a component literally called Harness Runtime: each agent session in its own Firecracker microVM, hosting Claude Code, Codex CLI, OpenCode or a custom agent, with a separate Action Gateway that brokers credentials and governed access to more than 16,000 tools. Renting the runtime is reasonable. But the runtime is the execution environment row of that table. The tool registry, the policy layer and the permission rules are still empty fields until somebody in your organization fills them in.
What the overhead buys, and who decided
The fair objection to “harness tax” came up within hours on Hacker News: a lot of Claude Code’s extra weight is permission checks, sandboxing, safety instructions and a tool surface that includes subagents and MCP servers. Pi leaves most of that out on purpose. Calling the difference a tax, the argument goes, treats security as dead weight.
That’s right, and it’s also the point. The overhead is a set of decisions — which tools are exposed, what the model is told about how to behave, what runs without asking, how much context is kept — that a vendor made once for every customer at once. In a coding shop that’s a productivity setting. In an enterprise agent platform it’s the control plane. It’s the layer that decides whether “read-only” is enforced on the request or on the effect, the layer that decides which parts of a workflow are allowed to be probabilistic, and the first thing a shadow-AI inventory has to find. Nobody in the buying organization reviewed it, because nobody knew it was there to review. They picked a model. The harness came in the box.
So the harness is three policies at once, and none of them has a name against it. A spend policy: 27,000 tokens of fixed context per call, times turns, times sessions, times seats, on a metered bill that already has no owner. A security policy: what the agent may touch and when it asks. An integration policy: which systems it can reach and through what schema. Change the version and all three change silently. The Berkeley tables pin the versions — Claude Code 2.1.224, Codex 0.146.0, Pi 0.85.1 — because the authors knew the numbers would drift with the next release. Your finance team’s dashboard doesn’t have a column for that.
What the working version looks like
Not a switch to the cheapest harness. A harness with an owner.
Start by measuring what you’re already sending. A gateway in front of the model endpoint, logging per request the instruction length, the tool definitions declared, and the provider’s input tokens split into uncached, cache-write and cache-read. Read the first main call of a session before anything else. That number is the fixed cost every later turn inherits, and it’s usually the one nobody in the room has seen. It’s also where the surprises live: a tool schema that triples because someone added description text, a system prompt that quietly doubles at a version bump, a cache prefix that stops matching after a plugin reorders the tool list.
Then write the harness down as configuration, the way the Berkeley study had to in order to run at all: the model, the allowed tools and their schemas, the instruction set, the context and turn limits, the sandbox and permission rules, the version. Put it in version control next to the code it serves. Diff it on every upgrade. When the tool count goes from 23 to 26, that’s a change request, not a release note.
ASD’s guidance closes with seven governance questions for directors. The last one is the test for whether your harness has an owner: if it were compromised, misconfigured or manipulated, what is the worst outcome, and what would limit it? If the room goes quiet, you’ve found the gap.
Then run the bake-off on your own work, not on SWE-bench. Thirty of your real tasks, the same model, two or three harnesses, three attempts each, one price list. Berkeley’s design is reproducible on a weekend. Thirty tasks is enough to see a 2× cost difference with confidence; it isn’t enough to see a 2% success difference, which is precisely why the cost finding is the durable one. If your workload needs the sandbox and the permission prompts, keep them and know what they cost. If it needs a legal index and a matter-file plugin, that’s your Astra, and it’s integration work with a schema, not a model upgrade.
Same model, twice the bill, and the answer didn’t change. Somebody should have been able to say why.
FAQ
- What is an AI agent harness?
- Everything in an agent system except the model. The Australian Signals Directorate's guidance 'Agentic AI harnesses: the layer above the model' (11 September 2026) uses the term for 'the components of an agentic AI system other than the LLM itself': the software that connects the model to tools, data sources, memory and planning workflows, enforces what it's allowed to do, and runs the loop until the task is done. ASD's shorthand is that if the LLM is the brain, the harness is the body. In practice that means the system instructions, the tool registry and its schemas, how context is assembled and trimmed, the permission system, the sandbox, memory, connectors and the audit log. Claude Code, Codex CLI and Pi are three harnesses for the same models, and UC Berkeley Sky Lab's HarnessTax study (16 September 2026) makes the point that choosing a coding agent means choosing a harness even when you think you're choosing a model. An Agentforce agent, a Copilot agent or an AgentCore deployment is the same harness decision, just presented as a model decision.
- AI agent harness vs model: what's the difference?
- The model predicts text from whatever is put in front of it. The harness decides what goes in front of it and what happens to the output. ASD's comparison table puts it plainly: the model is stateless, produces tokens only, and is 'a pluggable part you can swap'. The harness is stateful, executes tools, edits files and calls APIs, holds memory and session logs, and carries the permissions, sandboxes and approval gates. ASD calls it 'the durable architecture that outlives any single model'. The practical consequence is that cost, reliability and security risk are mostly decided in the harness, which is also the only part your organization configures.
- Who is responsible for securing an AI agent harness?
- You are, whether you built the harness or bought it. ASD's first key takeaway is that organisations control the harness, not the LLM, and its guidance says that for prompt injection 'no fully reliable technical mitigation currently exists' because models can't reliably tell instructions from the data they read. So the controls have to sit in the harness: what the agent can reach, which actions need human approval, what gets logged. ASD is explicit that no harness is inherently secure, and that this holds whether the harness was developed in-house or supplied inside a commercial product.
- Does the harness actually change how well a coding agent performs?
- On the two open benchmarks Berkeley tested, barely. HarnessTax ran 21 model–harness pairs (seven models across Claude Code, Codex CLI and Pi) on 30 sampled tasks each from SWE-bench Lite and Terminal-Bench 2.0, three attempts per task, with a 100-turn cap. The average harness effect on success rate stayed within ±2% on SWE-bench Lite and about ±5% on Terminal-Bench 2.0. Claude Fable 5 solved 97.8% of attempts in Claude Code, 96.7% in Codex and 96.7% in Pi. In nine of twelve comparisons across the six Anthropic and OpenAI models, a harness other than the model vendor's own produced the highest observed success rate. The authors' own caveat: two public benchmarks the models may have seen in training, and other workloads may differ.
- Why does the same model cost 2× to 5× more in one harness than another?
- Because the harness decides what is sent on every call before the task prompt is added. In Berkeley's measurements on SWE-bench Lite, Pi declares four tools with about 2,900 characters of tool schema and 2,500 characters of instructions, for a mean first-call context of roughly 1,970 tokens. Claude Code declares 23 tools with about 77,000 characters of schema and 13,500 characters of instructions, for a mean first call of roughly 27,000 tokens — over 10× Pi's, across all seven models. Portkey measured the same shape independently in April 2026 by routing the three agents through its gateway: about 2,600 input tokens per request for Pi, 15,000 for Codex, 27,000 for Claude Code. Multiply that by fifteen turns and the bill separates even when the answer doesn't. Caching softens it; cache reads are still billed, and Berkeley's first-call number already includes them.
- How do we measure our own harness overhead?
- Put a gateway or proxy in front of the model endpoint and log, per request, the instruction length, the number and size of tool definitions declared, and the provider-reported input tokens split into uncached, cache-write and cache-read. Read the first main call of a session before you read anything else, because that's the fixed cost every turn inherits. Then run your own bake-off the way Berkeley did: a fixed set of your real tasks, the same model, two or three harnesses, three attempts each, the same price list. Thirty tasks is enough to see a 2× cost gap; it is not enough to see a 2% success gap, and that asymmetry is the point.
- Should we just switch to a minimal harness?
- Not on the strength of a benchmark. The overhead in a vendor harness buys real things — permission prompts, sandboxing, subagents, a larger tool surface — and a minimal harness like Pi omits most of them by design. The operator problem isn't that Claude Code costs more than Pi; it's that in most organizations no one decided what the extra context is buying, no one can name which of the 23 tools the workload uses, and the harness version isn't recorded anywhere as a change. Write the harness configuration down as a policy with an owner — allowed tools, instruction set, context and turn limits, sandbox rules — and the switching question answers itself per workload.
- What does OpenAI's Astra for Law have to do with harness cost?
- It's the same finding from the other direction. Astra for Law, announced 17 September 2026, is GPT-6 Astra plus a legal search index, custom instructions for legal analysis and writing, and 26 partner plugins into tools like iManage, Intapp, Relativity and Clio. On 200 questions from Vals AI's Legal Research Bench, OpenAI reports the configured system passed the correctness check on 54.0% of questions versus 38.7% for the same model with web search alone — a 40% relative gain from the harness, with the model held constant. Berkeley showed the harness can move cost without moving success; OpenAI showed it can move success when it carries domain-specific tools and instructions. Either way, the harness is the variable, and it's the one most buyers never evaluate.
- Can you rent an AI agent harness as a managed service?
- The runtime, yes. DigitalOcean's Managed Agents, in public preview and covered by InfoQ on 2 October 2026, sells a component literally named Harness Runtime: each agent session runs in its own Firecracker microVM, and it hosts existing harnesses such as Claude Code, Codex CLI and OpenCode. A companion Action Gateway brokers credentials and governed access to more than 16,000 tools, with optional human approval for sensitive operations. What you can't rent is the configuration: which tools a given workload is allowed, what its instructions say, which actions need a person. Hosting moves the compute. The policy is still yours to write.