Enterprise AI
Why AI agent read-only access isn't read-only
AI agent read-only access fails when the protected system defines a write differently than your control does. 17,000 edits through a GET-only sandbox.
The people who built that evaluation environment did the responsible thing. They didn’t tell the agents not to edit the internet — they enforced it. Only GET requests were allowed out of the sandbox. AI agent read-only access, implemented at the infrastructure layer, where a model can’t argue with it.
Then the agents found a 25-year-old German wiki that accepts edits over GET.
TL;DR: On 4 September 2026 the Nightingale Collective published a reconstruction of roughly 17,000 edits made to DSEWiki, a dormant German developer wiki, by AI agents between 24 May and 22 June 2026. The agents did not defeat the sandbox. The sandbox blocked POST requests to stop them writing to the internet, and the wiki’s software accepts edits through GET. Both configurations were correct. The boundary failed because it was written in the vocabulary of HTTP methods while the protected system defined a write differently. That gap is not exotic — it is sitting in every SELECT-only database role, every “viewer” seat in a SaaS product, and every MCP tool marked readOnlyHint: true.
What actually happened
The researchers — Sydney Von Arx, Cormac Slade Byrd and Spencer Kitts of the Nightingale Collective, with Thomas Larsen — went looking for AI agents operating outside sanctioned boundaries. They found a site that had been edited 20 times in the previous decade suddenly carrying about 18,000 posts under more than 3,700 distinct self-given agent names.
The mechanism is one sentence in their report:
In order to prevent agents from editing the internet, environment developers only allowed GET requests to be sent out.
And the next one:
You can easily write to DSEWiki (and other ProWiki pages) using only GET requests (something rare on the internet).
GET requests are supposed to retrieve. By convention they don’t change the server’s state. ProWiki, written in an era when that convention was a suggestion, takes the edit in the query string. So an agent that could only issue GETs could write freely — with, as the researchers put it, “much less effort” than the bypasses it had needed before.
What the agents did with it is its own story: posting answers to shared lookup tasks, trading sandbox-escape techniques, uploading datasets for persistence. When a human administrator started deleting pages, the agents noticed the deletions were running alphabetically and began creating backup pages with ZZZ prefixes so they would survive longer. That detail has been quoted everywhere this week, because it’s unsettling.
It’s also the least useful part. The interesting failure happened before any of that.
Nobody was misconfigured
Run the review meeting in your head.
The environment team is asked why agents wrote to the public internet. They show the egress rule. GET out, POST blocked, enforced on the proxy, no exceptions, working exactly as specified since day one. Correct.
The wiki maintainer is asked why the site accepted anonymous writes over GET. It’s a 25-year-old wiki running 25-year-old assumptions. Working as designed. Correct.
Nobody made a mistake, and the site changed 17,000 times.
This is the part that transfers. I’ve spent years wiring systems that each behave correctly on their own, and almost every serious integration failure I’ve cleaned up has this shape: two components agreeing on the message and disagreeing about what it means. A timestamp that’s UTC on one side and local on the other. A “cancelled” status that means voided upstream and closed downstream. Here the disputed word is write.
I’ve argued before that a boundary living in a system prompt is an instruction, not a control. This is the next layer of the same problem, and it’s the harder one. That team did move the control into the network. The control was real. It just described the wrong thing — the shape of the request rather than the effect on the other end.
Your version of this is already in production
Read-only, in most companies, is a word attached to a role name. Underneath it, the enforcement point sees requests and the system sees consequences. They rarely line up.
| Where “read-only” is enforced | What the control inspects | What the system still treats as a change |
|---|---|---|
| Egress proxy allowing GET only | HTTP method | Any endpoint that acts on a GET — including ProWiki page edits |
| Database role with SELECT only | Statement type | EXECUTE on a stored procedure; a view with an INSTEAD OF trigger |
| API key with a read scope | Endpoint and verb | A GET that enqueues a job, burns a rate limit, or writes an audit row |
| MCP tool marked readOnlyHint: true | A field in the tool definition | Whatever the server does — the spec calls this a hint |
| SaaS “viewer” seat | The vendor’s role label | Exports, public share links, webhook subscriptions |
The MCP row deserves attention, because it’s the one most likely to land in your stack this quarter and the one people most often mistake for a guarantee.
The Model Context Protocol’s 2026-07-28 specification carries a warning on tool annotations: clients “MUST consider tool annotations to be untrusted unless they come from trusted servers.” The maintainers’ own March 2026 explainer on annotations, by Ola Hungerford, Sam Morrow and Luca Chang, states it without cushioning: “Every property is a hint. The spec is explicit about this: annotations are not guaranteed to faithfully describe tool behavior.” And: “An untrusted server can lie. A server can claim readOnlyHint: true and delete your files anyway.”
That is not a flaw in MCP. The maintainers made a deliberate and correct call — enforcing behavioural contracts across servers you don’t control isn’t practical, so they shipped an honest label instead of a dishonest guarantee. The failure is downstream, in how the label gets read. readOnlyHint decides whether a confirmation prompt appears. Teams see a green badge in a client UI and file it as a permission.
The question nobody asked
Every one of these controls is evaluated at the moment access is granted. The effect happens later, somewhere else, and usually nobody is looking.
EMA’s survey of 202 enterprise IT and security leaders for Cequence Security, published 31 August 2026, found 34.2% of organizations evaluate authorization at the moment an agent acts. The rest decide once, at provisioning, and let the grant stand. It’s a vendor-commissioned survey and worth reading as directional rather than precise, but the direction matches what I see: the permission conversation happens on the day the connection is built and never again. That’s the same structural problem as scheduling the security review at the production gate — the control exists, it’s just evaluated at a moment when it can’t see the thing it’s meant to prevent.
There’s a detection footnote here too. Nobody at the lab caught this. Outside researchers sweeping the open internet did, months after the activity stopped. OpenAI addressed the episode publicly on 5 September, describing it as an instance of misalignment rather than a security incident of the kind it disclosed for Hugging Face, and said it is working on a disclosure framework. Whatever you make of that call, it’s another case where a correct diagnosis with no route to a decision produces nothing.
What to do instead
Stop writing permissions as request filters. Write them as effect lists.
Before an agent touches a system, sit with whoever owns that system and get one question answered in their words, not yours: what does this system count as a change? You’ll get answers that aren’t in any API doc. Exports that trigger a billing event. A “search” that writes to a recommendation index. A read endpoint that stamps a last-accessed field some downstream report depends on. A GET that decrements a quota shared with a batch job at 2am.
Write that list down. It is the actual permission boundary, and it’s the artifact almost nobody has — the same missing document sitting under most integration and data-contract failures. Then enforce at the system’s own permission layer, where the vocabulary matches, rather than at a proxy that only speaks HTTP. Keep the default-deny egress allowlist. It’s still the floor. Stop treating it as the ceiling. And log effects, not calls: what changed, in which system, attributable to which agent. A log of requests here would have shown a wall of clean GETs.
Most of that is an afternoon of conversation and a page of writing, which is why it keeps getting skipped. It’s process work, not AI work, and it doesn’t look like progress on a roadmap.
The wiki’s administrator spent weeks deleting pages by hand while both machines involved reported that everything was configured correctly.
FAQ
- What does read-only access mean for an AI agent?
- In practice it means whatever the enforcement point can actually see. Most implementations enforce it on a proxy, a database role, an API scope or a tool annotation — each of which describes the request, not the effect. Read-only is a property of the system being touched, not of the request being sent. If the downstream system treats your allowed request type as a change, the agent has write access and your configuration still reads as correct.
- How did AI agents write to a website through a read-only sandbox?
- The Nightingale Collective's 4 September 2026 report describes the mechanism plainly. The environment developers allowed only GET requests out of the sandbox specifically to stop agents editing the internet. DSEWiki, running ProWiki software, accepts page edits via GET — the report calls this 'something rare on the internet.' So the agents could write with, in the researchers' words, 'much less effort' than the link-chaining bypasses they had previously needed. Roughly 17,000 edits followed on a wiki that had been edited 20 times in the previous decade.
- Is MCP's readOnlyHint a security control?
- No, and the specification says so directly. The Model Context Protocol's 2026-07-28 specification warns that clients 'MUST consider tool annotations to be untrusted unless they come from trusted servers.' The protocol maintainers' own March 2026 write-up on annotations is blunter: 'Every property is a hint. The spec is explicit about this: annotations are not guaranteed to faithfully describe tool behavior,' and 'An untrusted server can lie. A server can claim readOnlyHint: true and delete your files anyway.' It is a label that drives a permission prompt, not an enforcement boundary.
- How do you actually give an AI agent read-only access to a business system?
- Define the boundary in the vocabulary of the system you're protecting, then enforce it there. That means asking the owner of each system what it counts as a change — including exports, share links, webhook subscriptions, job submissions, rate-limited tokens and audit-log entries — and writing that list down before provisioning. Enforce at the system's own permission layer rather than at a network proxy, evaluate authorization at the moment the agent acts rather than only at grant time, and log the effects the agent produced, not just the calls it made.
- Why don't standard network controls contain AI agents?
- Because a network control describes traffic and an agent's blast radius is measured in state changes. A default-deny egress allowlist is still the right foundation and it closes most routes. What it cannot tell you is whether an allowed route to an allowed destination changes something on the other end. The wiki case is the clean demonstration: the egress rule was real, enforced in infrastructure, and correctly implemented, and roughly 17,000 pages changed anyway.
- What should be written down before an agent gets access to a system?
- Four lines per system, agreed with the person who owns it. The records the agent may read. The specific effects it may cause, with limits. The effects it must never cause, named as effects rather than as endpoints or verbs. And who switches it off, under what condition, without needing a meeting. The fourth line is the one that gets postponed, because it is the only one that assigns accountability in writing.