Aviation
AI crew scheduling: the redundancy is at the wrong layer
AI crew scheduling now gets dual-cloud failover. The worst crew collapse in Ryanair's history involved nothing going down — a calendar definition changed and the plans underneath it didn't.
Ryanair signed a five-year data and AI deal with Google Cloud on 12 August 2026, putting Gemini Enterprise to work on flight crew logistics and decision automation. The reason its chief executive gave was resilience — a dual-cloud setup so critical services keep running if one provider has problems. So AI crew scheduling at an airline carrying around 216 million passengers a year now runs on infrastructure with a spare.
The worst crew scheduling failure in that airline’s history involved no infrastructure at all. Nothing went down. A calendar definition changed, and the plans built on top of it didn’t change with it.
TL;DR: Dual-cloud failover protects the layer that rarely breaks. In September 2017 Ryanair cancelled 2,100 flights and pulled 25 aircraft out of its winter programme because it moved its holiday leave year from April–March to a calendar year and had to compress the transition into nine months. Every system stayed available the whole time. The definition of a leave year changed, the plans that consumed it weren’t recomputed against the new window in time, and there is no failover for that. AI crew scheduling pays off once the definitions are settled — ANA’s took roughly four years of proof-of-concept before it went into full operation.
What Ryanair actually bought
The announcement is specific, which makes it useful. Per Google Cloud’s 12 August 2026 release, the five-year agreement puts Google Workspace and Google Cloud in front of 35,000 employees. Gemini Enterprise, described as Google’s agentic AI platform, is aimed at automating decision-making, optimising flight crew logistics, and corporate productivity. DeepMind’s AlphaEvolve and WeatherNext are pointed at fleet operations and maintenance scheduling. The framing is growth: roughly 216 million passengers a year today, 300 million targeted by 2034.
The resilience piece, in Google Cloud’s own words, is about building a flexible system so Ryanair can use multiple cloud providers — if one platform has issues, critical airline services adapt quickly and keep running.
CEO Eddie Wilson tied the deal to exactly that. Ryanair is on a growth journey to 300 million passengers by 2034, he said, and to support it the airline needs excellent infrastructure resilience, which the new dual-cloud strategy provides.
I have no quarrel with any of it. Dual-cloud is sound engineering. Gemini on crew logistics is a reasonable place to put an agent, and aiming an evolutionary optimizer at fleet operations is better-targeted than most enterprise AI announcements manage. The question is narrower: resilience against what?
The failure that had no outage
Ryanair published the whole thing itself, which is rare enough to take at face value.
On 15 September 2017 the airline announced it would cancel 40 to 50 flights a day for six weeks — under 2% of its 2,500 daily flights — after punctuality dropped from 90% to under 80% over two weeks. It listed the causes: ATC capacity delays, strikes, weather, and the impact of increased holiday allocations to pilots and cabin crew.
That last one is the structural cause, and the mechanism sits in the same statement. Ryanair had agreed with the Irish Aviation Authority to run a nine-month annual leave transition period from April to December 2017, moving from a holiday leave year of April 2016 to March 2017 over to a calendar leave year running January to December 2018. Leave had to be allocated before 31 December 2017 to start the new calendar year with no backlog.
Read that as an operator and the arithmetic is ugly. A twelve-month entitlement has to be consumed inside a nine-month window. Crew availability in the back half of 2017 was therefore structurally lower than in any normal year — not because anyone was sick, not because of weather, but because the definition of the year they were counted against had moved. The flying programme underneath was still built for a normal year.
The 27 September statement confirms both the damage and the fix: 2,100 of over 800,000 annual flights cancelled, and the airline flying 25 fewer aircraft of its 400 from November 2017, then 10 fewer of 445 from April 2018, specifically to roster the extra leave into October, November and December and hit the IAA’s deadline.
Michael O’Leary’s line at the time was that from that day there would be no more rostering-related cancellations that winter or in summer 2018. The remedy was to shrink the operation until the plan matched the constraint. That is what it costs to reconcile after the fact.
Not one minute of that was an availability problem. Every system was up. Every number in every system was correct under its own rules. You can fail over between clouds. You can’t fail over between two definitions of a leave year.
The field that’s right in both systems
This failure mode doesn’t look like a failure until it’s expensive, and it isn’t specific to airlines.
Somewhere in a crew system there’s a field that means remaining annual leave. Somewhere in an HR system there’s a field with the same name. Both are computed correctly. Both are computed over a window, and the window is a definition that lives outside both systems — in a policy document, in a regulator’s requirement, in a decision made in a meeting. Change the window and neither system throws an error. They start producing numbers that can’t both be true of the same plan, and nothing in the stack is designed to notice.
Monitoring won’t catch it, because monitoring watches for things that stop responding. The capacity plan didn’t stop responding. It answered confidently, using last year’s window.
That’s the same defect I’ve written about in integration and data contracts: the seam between two systems is where meaning is supposed to be agreed, and it’s the part nobody owns. A leave year is a data contract. It has a definition, consumers, and a change process. Ryanair’s changed — for a legitimate regulatory reason — and the downstream recomputation didn’t happen fast enough.
Adding a second cloud does not add a second opinion about what a leave year is. It adds a second copy of the same wrong assumption, in a different region, highly available.
What four years of ANA says about the real work
If you want to know where the effort actually goes in AI crew scheduling, the most instructive announcement of the summer wasn’t Ryanair’s.
On 21 July 2026, ANA Holdings and R&D — a University of Tokyo startup — announced that their AI Scheduler, abbreviated AIS, had entered full operation in April 2026 for roughly 2,000 flight crew. It followed approximately four years of proof-of-concept work. The system handles crew qualifications, aviation regulations, internal company rules, fatigue risk management standards, and individual leave requests, which the companies describe as tens of thousands of constraint combinations.
Look at that list again. Individual leave requests are in it — as a constraint the solver has to satisfy, precisely defined, sitting alongside the regulations. The exact category that took 2,100 Ryanair flights off the board is, in ANA’s system, a modelled input.
The published payoff is 6,300 hours a year, 525 a month. Across about 2,000 crew that works out to a little over three hours per crew member per year of scheduling work removed — my arithmetic on their two published figures, not a claim they make. It’s a genuine result and I’d take it. It’s also unmistakably a back-office number, produced by four years of work that was mostly not about the model.
Four years doesn’t go into training anything. It goes into getting every rule written down precisely enough, and agreed across enough departments, that a solver’s output can be trusted onto a live roster. The same pattern shows up in field service, where the constraints the template can’t express decide whether a schedule is usable, and in the utility scheduler I built, where the rules had never existed on paper before the project started.
| Ryanair 2017 | ANA 2026 | Ryanair 2026 | |
|---|---|---|---|
| What changed | Leave year moved to a calendar year over a 9-month transition | AI Scheduler into full operation after ~4 years of PoC | Five-year Google Cloud deal; Gemini Enterprise on crew logistics |
| Layer addressed | None — the definition changed, the plans didn’t | The constraint model: quals, regs, internal rules, FRMS, leave | Infrastructure (dual-cloud) plus an agentic layer |
| Measured outcome | 2,100 flights cancelled; 25 aircraft cut from Nov, 10 from Apr 2018 | 6,300 hours/year of scheduling work removed, ~2,000 crew | Announced 12 Aug 2026; no operational results yet |
| Time to result | Failure surfaced inside one transition window | ~4 years | Five-year term |
Sources: Ryanair corporate statements of 15 and 27 September 2017; the ANA Holdings and R&D announcement of 21 July 2026; Google Cloud’s release of 12 August 2026.
Redundancy at the layer that breaks
Failover is the right instinct pointed at the wrong object. The instinct says: this thing is critical, so keep a spare. Correct. The mistake is deciding the critical thing is the compute.
For any shared definition your operation depends on — leave year, duty period, base assignment, qualification currency, service window — the equivalent of a spare is three things no cloud contract can give you:
- A named owner, the only person allowed to change it, accountable for what happens when they do.
- A list of every system and plan that consumes it, so a change has a blast radius you can read before you make it rather than after.
- A reconciliation that runs on change and fails loudly when two consumers disagree — the alert Ryanair’s stack couldn’t have raised, because nothing was configured to compare the entitlement window against the roster capacity window.
None of that is exotic. All of it is boring. It is also the only thing that would have protected the winter 2017 schedule, and none of it arrives with a cloud migration or an agent platform.
The agentic layer Ryanair is deploying will be good at what agents are good at — reading a disruption, drafting the crew notification, answering a pilot’s question about their own roster. Real time savings. They sit downstream of the definitions, which means the agent inherits whatever disagreement is already there and applies it faster. An agent working from an entitlement calculated over the old window produces confident, well-written, wrong answers at scale.
The model picks up the phone. The definitions still have to agree.
If your scheduling problem looks like this — the systems are up, the numbers are right, and the plans still don’t fit — that’s the work I do. See how I approach it, or get in touch.
FAQ
- What is AI crew scheduling?
- It's the use of optimization and, more recently, language models to build the roster that assigns pilots and cabin crew to flights. The scheduling itself is a hard-constrained allocation problem: every assignment has to satisfy licence and type qualifications, flight-time and duty limits, base and standby rules, fatigue risk standards, union terms, and accrued leave. ANA's system, built with the University of Tokyo startup R&D and in full operation since April 2026, is described in the companies' 21 July 2026 announcement as handling tens of thousands of constraint combinations across roughly 2,000 flight crew. That's a solver's job. The newer agentic layer sits around it — reading disruption notices, drafting crew communications, explaining why a roster changed, handling the exception conversation. Both depend on the same thing: that the rules and the entitlements they consume are defined identically everywhere they're read.
- Can AI agents do airline crew rostering?
- Not the assignment itself, and not because the models aren't good enough. Rostering is combinatorial optimization, which solvers handle well and language models handle badly — a model that is fluent about duty limits is not the same as a system that provably never violates one. Where agents genuinely earn their place is at the edges of the solve: interpreting an irregular-operations event, drafting the notification, answering a crew member's question about their own roster, summarizing why the schedule changed overnight. Ryanair's August 2026 announcement with Google Cloud reflects that split — Gemini Enterprise for crew logistics and decision automation, and DeepMind's AlphaEvolve and WeatherNext pointed at fleet operations and maintenance scheduling, which are optimization and forecasting problems rather than conversation.
- What caused Ryanair's 2017 flight cancellations?
- Ryanair's own statements name it directly. On 15 September 2017 the airline said it would cancel 40 to 50 flights a day for six weeks — under 2% of its 2,500 daily flights — after punctuality fell from 90% to under 80% in two weeks. It listed ATC capacity delays, strikes, weather, and the effect of increased holiday allocations to pilots and cabin crew. That last item was the structural one. Ryanair had agreed with the Irish Aviation Authority to run a nine-month annual leave transition period from April to December 2017, moving from a holiday leave year of April to March to a calendar leave year from January 2018. Its 27 September statement puts the total at 2,100 of over 800,000 annual flights and confirms the fix: flying 25 fewer aircraft of its 400 from November, and 10 fewer of 445 from April 2018, to create enough roster capacity to clear the leave backlog before 31 December.
- Does dual-cloud architecture protect against operational failures?
- It protects against one class of failure — a provider or region becoming unavailable — and that class is worth protecting against. It does nothing for the class that actually causes most operational collapses, where every system is up and returning answers, and the answers disagree because two systems computed them against different definitions. Failover is a mechanism for substituting an identical thing when one copy stops responding. When the problem is that two copies are both responding and both correct under their own rules, there is nothing to fail over to. The redundancy you need there isn't a second cloud. It's a single owner for the definition and a reconciliation that runs whenever it changes.
- What is a data contract in crew scheduling?
- A data contract is the written, owned agreement about what a shared field means, who computes it, over what window, and what happens downstream when the rule changes. In crew scheduling the obvious candidates are the leave year, accrued and remaining entitlement, duty period boundaries, base assignment, and qualification currency. Each one is read by several systems — the roster optimizer, the HR system, the capacity plan, the crew app — and each of those can hold a technically correct value that disagrees with the others. The contract is what makes disagreement detectable. Without one, the systems don't argue; they quietly produce plans that can't both be executed, and you find out at the gate.
- How long does it take to deploy AI crew scheduling?
- Longer than the announcement cycle suggests. ANA and R&D ran approximately four years of proof-of-concept work before the AI Scheduler entered full operation in April 2026, and the published payoff is 6,300 hours a year — 525 hours a month — of scheduling work removed across roughly 2,000 flight crew. That is a real, defensible number, and it is a back-office number. Four years didn't go into the model. It went into encoding qualifications, aviation regulations, internal rules, fatigue risk standards, and individual leave requests precisely enough that a solver could be trusted with the output. The rule-writing is the project. The model is what you add once the rules exist.