Enterprise AI
Long-running AI agent tasks break on the wait, not the work
Long-running AI agent tasks fail in the gaps between steps. MCP's 22 August 2026 roadmap and AWS's session defaults show exactly where the state goes missing.
The maintainers of the Model Context Protocol published a new roadmap on 22 August 2026, and the first item on it is a quiet admission about long-running AI agent tasks: the request-and-response shape the whole ecosystem was built on doesn’t fit the work. Loops run longer than a call. Results come back after everyone has stopped listening, and somebody needs to be able to steer the thing while it’s still moving. That’s a protocol-level problem now, which means enough production deployments have hit it that it stopped being anyone’s local bug.
TL;DR: Long-running AI agent tasks don’t fail because the model runs out of stamina. They fail because real processes spend most of their duration waiting on someone else, and nothing in a standard agent deployment owns the state during that wait. Amazon Bedrock AgentCore Runtime ships with a fifteen-minute idle timeout and an eight-hour maximum lifetime that, per AWS’s docs, cannot be reset. Your permit takes eleven days. The gap between those two numbers is a process specification nobody wrote.
The protocol is being rebuilt around the gaps
The 22 August roadmap, published by lead maintainers David Soria Parra and Den Delimarsky, lists five priority areas. Agentic messaging primitives comes first, ahead of HTTP-native transport, agent identity, improved primitives, and SDK developer experience.
The reasoning is stated plainly: “Modern agentic workloads no longer fit the standard request-and-response pattern. Loops can run for longer, servers can push streamed results, and there is a clear need to steer work mid-flight.”
What’s being built in response is a list of things that only matter once work outlives the call. Server-initiated events — webhooks and channels, described in the roadmap as being there “so clients aren’t left polling for results.” The Tasks extension, SEP-2663, matured far enough to move into the specification proper. Composition work spanning three separate working groups, Agents, Transports, and Triggers & Events, because none of them can solve it alone.
This isn’t the first signal. MCP’s March 2026 roadmap already noted that early production use had surfaced lifecycle gaps in the Tasks primitive: retry semantics when a task fails transiently, and expiry policies for how long results are retained after completion. Five months later the async problem is the top line.
The identity section of the same roadmap contains the other half of the story, and it’s worth reading as an operator. The stated goal includes “helping developers avoid pasted API keys and long-lived tokens,” with work on Demonstrating Proof of Possession, Workload Identity Federation, and standard token exchange, plus engagement with the IETF OAuth and WIMSE working groups. Read that sentence again. The people who maintain the protocol are describing the current state of enterprise agent authentication as pasted keys and long-lived tokens. That’s not a criticism of anyone’s engineering. It’s what happens when a thing designed for a session gets used for a process.
Duration is not the same as work
The industry measures agent capability in how long the agent can work. That is the wrong axis for almost every business process, and the mismatch is the whole problem.
METR defines its task-completion time horizon as “the task duration (measured by human expert completion time) at which an AI agent is predicted to succeed with a given level of reliability.” It’s a careful, useful measurement, and METR is unusually honest about its limits: the task distribution is “limited primarily to software engineering, machine learning, or cybersecurity tasks,” the tasks are “much ‘cleaner’ than real economically valuable labor,” and “AI agents did worse on messier tasks.” The page also notes that measurements above 16 hours are unreliable with the current suite.
Now put that next to an actual operation.
A change order sits with the architect for six days. A prior authorization runs against a payer clock measured in business days. A quote waits on a supplier who answers on Thursdays. A tenant maintenance request waits on a part. A background check comes back when it comes back. In every one of those, the work is maybe forty minutes. The duration is a week and a half, and 99% of it is a queue in somebody else’s building.
Doubling an agent’s work horizon does not shorten a six-day wait. The horizon was never the constraint. What the wait actually demands is something nobody benchmarks: the ability to stop, hold everything relevant, survive whatever happens to the runtime in the meantime, and pick up with the same authority and the same context when an answer finally lands.
That’s a state management problem wearing an AI costume.
Fifteen minutes
Here’s the number that makes this concrete, straight out of the AWS documentation for Bedrock AgentCore Runtime lifecycle settings.
| Setting | Default | Maximum | Behavior |
|---|---|---|---|
idleRuntimeSessionTimeout | 900 seconds (15 minutes) | 28,800 s on microVMs; 1,209,600 s on Instances | Resets each time you invoke the same session |
maxLifetime | 28,800 seconds (8 hours) | 28,800 s on microVMs; 1,209,600 s on Instances | Starts at microVM creation and cannot be reset |
Fifteen minutes idle, eight hours total, out of the box. AWS added an Instances compute type in August 2026 that raises the ceiling to persistent sessions of up to 14 days — a genuine improvement, and still shorter than an ordinary procurement cycle.
I’m not picking on AWS here. These are sane defaults for the thing they’re defaults for, they’re documented clearly, and AWS’s own guidance says to start with them and adjust. The point is what happens when nobody does the adjusting, because nobody wrote down how long the process takes. The agent that was going to handle your permit application gets terminated at minute fifteen of a ten-day wait, and the termination is a normal, successful, unremarkable event in the logs. There is no error. There’s a session that isn’t there anymore.
Then the response arrives. The webhook fires into a session that ended eight days ago. If nothing durable on your side was listening, the process doesn’t fail — it just stops, and it stops in the silent way, which is the way that takes three weeks to notice. This is the same failure shape as an agent dependency chain nobody monitors end to end: every component reports healthy, and the work is gone.
What actually goes missing during the wait
Five things have to survive a pause, and in most deployments none of them has a named owner.
Progress. What the agent had done and how far it got, recorded somewhere other than a model context window. If resumption means replaying the conversation, you’re storing your process state in a transcript — a pattern with its own retention problem.
Authority. What the agent is permitted to do when it wakes up. Credentials expire, approvals lapse, the requester changes roles, the budget period closes. An agent that resumes with the permissions it had eleven days ago is exercising an authorization nobody re-granted.
The completion condition. What counts as the answer it was waiting for. This is where it gets genuinely hard, and it’s hard for reasons that have nothing to do with AI. Two responses arrive and disagree. The reply comes as a PDF attachment on an email thread nobody routed. Someone answers verbally in a meeting and the system never hears about it. If a human currently resolves that ambiguity by knowing who to trust, you have an undocumented rule, and the agent will not infer it.
The deadline. How long the wait is allowed to last before somebody is told. A timeout that logs is not an escalation. Escalation names a person — and naming one is a separate question from whether that person catches anything, which has its own measured yield.
Ownership. Who holds the work while the agent is dormant. The honest answer in most deployments is nobody, which is exactly how a request spends a fortnight in a state no dashboard displays.
That list is a specification of your process. It’s not a platform capability, which is why extending a session ceiling from 8 hours to 14 days doesn’t produce it. Longer sessions buy time for work that was already defined. They don’t define anything.
Write the wait down first
The practical version is short, and it’s the same discipline that makes multi-agent orchestration work or not work.
Take the process you want to automate and mark every point where it stops and waits on something outside your control. For each one, write five lines: what triggers the wait, who is accountable while it lasts, the maximum acceptable duration, what happens when that expires, and what the agent’s authority looks like when it resumes. Most processes have three or four such pauses. Some have eleven, and finding that out is usually the most valuable hour of the project.
Then look at what you’ve produced, because it will tell you something uncomfortable and useful. Frequently the biggest win isn’t the agent at all — it’s that you’ve just documented a six-day queue that existed because a form got emailed instead of submitted. Fixing that removes a step. The agent removes keystrokes. ROI shows up when you remove the step.
And if you do build it, build the wait as a first-class object: durable, owned, with a deadline and an escalation path, in a system that doesn’t disappear on an idle timer. The protocol maintainers are working toward primitives that make this cleaner, and that work is welcome. Tasks, webhooks, channels, mid-flight steering — all of it helps. None of it decides how long your subcontractor gets before someone picks up the phone.
That decision was always yours. It just never got written down, because when a human owned the process they simply knew. The agent doesn’t, and it will hold the request patiently forever, right up until the microVM terminates at minute fifteen.
FAQ
- Why do long-running AI agent tasks fail in production?
- Because the duration of a real business process is mostly waiting, not working, and the agent stack was built around a request that gets an answer. A permit takes eleven days because it sits in someone else's queue for ten of them. A quote waits on a supplier who answers Thursday. During that gap the agent isn't computing — it's supposed to be dormant, holding state, and ready to resume with the same authority it had when it stopped. Nothing in a standard request-and-response deployment owns that. MCP's lead maintainers made this the first priority of their 22 August 2026 roadmap, writing that "modern agentic workloads no longer fit the standard request-and-response pattern" and that "loops can run for longer, servers can push streamed results, and there is a clear need to steer work mid-flight."
- How long can an AI agent session actually run?
- Less time than most business processes take, unless someone changes a setting. Amazon Bedrock AgentCore Runtime ships with an idle session timeout of 900 seconds — fifteen minutes — and a maximum lifetime of 28,800 seconds, or eight hours. The idle timer resets each time you invoke the same session, but AWS's own documentation states the maximum-lifetime timer "starts when the microVM is first created and cannot be reset." AWS added an Instances compute type in August 2026 that supports persistent sessions of up to 14 days. Fourteen days is a real improvement and still shorter than plenty of ordinary approval cycles.
- What is the difference between a long-running task and a long-horizon task?
- Long-horizon usually means the agent works continuously for a long time. Long-running, in an operations sense, means the process spans a long time while the agent does almost nothing for most of it. METR measures the first: it defines its time horizon as "the task duration (measured by human expert completion time) at which an AI agent is predicted to succeed with a given level of reliability." That's a measure of work, benchmarked on software, ML and cybersecurity tasks — and METR is explicit that its "tasks are much 'cleaner' than real economically valuable labor" and that "AI agents did worse on messier tasks." A three-day wait for a subcontractor to send a certificate of insurance contains no work at all. Better endurance does not shorten it.
- What state does an AI agent need to survive a multi-day wait?
- Five things, and most teams have written down none of them. What the agent was doing and how far it got, in a form something other than the model can read. What it is allowed to do when it resumes, since credentials and approvals expire on their own schedule. What counts as the answer it was waiting for, and what to do if two arrive or none does. How long the wait is allowed to last before somebody is told. And who owns the work in the meantime — a named human, not the agent. That list is a process specification, not a platform feature, which is why buying a longer session doesn't produce it.
- Should I use webhooks or polling for long-running agent workflows?
- Webhooks, with a durable record on your side, and the protocol is moving that way. The 22 August 2026 MCP roadmap describes work on "server-initiated events (webhooks and channels, so clients aren't left polling for results)," alongside maturing the Tasks extension, SEP-2663, so it can move into the specification. But the transport is the easy half. A webhook that fires into a session that timed out four hours ago is a silent failure, and it fails quietly — no error, no alert, just a process that stops. If the callback has nowhere durable to land, the delivery mechanism doesn't matter.
- How do I make an AI agent handle approvals and human handoffs correctly?
- Write the wait down before you build anything. For each pause, record what triggers it, who is responsible while it lasts, what the maximum duration is, what happens when that expires, and what the agent's authority looks like on the other side. Then treat the deadline as a real event with an owner rather than a timeout that logs and moves on. Note that this is a separate question from whether the human catches errors when they do respond — approval as a review step has its own measured yield, which I've written about separately. Both problems have to be solved, and the duration one usually isn't even named.