If Your Automated Process Failed Safely Tomorrow, Could Your People Still Run the Business at Volume?
Reading the UK air traffic control chaos as a case study in the automation paradox, and the specific corporate governance question it now forces.

Sign up for my Substack daily AI newsletter here.
See my AI Training course portfolio for corporate Business Leaders here.
Follow me on LinkedIn: https://www.linkedin.com/in/johanosteyn/
Recently, more than 1,300 flights were cancelled across the United Kingdom over two days after a fault in the National Air Traffic Services flight data system forced controllers to revert to manual methods. The dominant reading in South African commentary, including a Citizen editorial published the following day, framed the incident as a warning about the risks of artificial intelligence in critical infrastructure and asked how protected our own systems are from “a digital bot going rogue.” That framing is understandable, and it is aimed at the wrong question. The NATS system did not go rogue. It failed exactly the way safety-critical software is supposed to fail. The problem was not the automation. The problem was what happened to the humans who had to take over when the automation stopped, and the question that arises for your organisation has almost nothing to do with AI.
CONTEXT AND BACKGROUND
To understand what actually happened at NATS, it helps to look at the previous incident of the same class. In August 2023, a nearly identical failure grounded another 1,500 UK flights. The Civil Aviation Authority’s preliminary report on that incident is unusually forensic and public, and its findings are worth reading before any commentary on last week’s chaos. A flight plan submitted through the European system contained two waypoints with the same identifier, roughly 4,000 nautical miles apart. The NATS flight-planning software correctly identified this as ambiguous data that could not be safely passed to controllers, raised what the report calls a critical exception, and stopped processing. The backup system, running the same software on separate hardware, applied the same logic and stopped too. In under 20 seconds, both the primary and backup systems were offline. Controllers reverted to manual entry of flight plans and had to reduce UK traffic flow drastically to stay within the capacity of manual handling. Martin Rolfe, the chief executive of NATS, ruled out a cyberattack within hours, and no evidence has emerged that AI played any role in either failure.
The specific mechanism matters. The NATS software did not malfunction. It behaved exactly as safety-critical software is designed to behave. When faced with a flight plan it could not resolve unambiguously, it faced a specific choice. It could have passed ambiguous data forward to controllers, who would then be routing aircraft based on flight plans the software knew might be wrong. Or it could stop. It stopped. The consequence of stopping was that traffic dropped to what human controllers could handle manually, and 1,500 flights were cancelled. The consequence of passing the data forward could have been aircraft on collision paths in controlled airspace.
When safety engineering succeeds in critical systems, operational disruption is often the price. The 2023 NATS software chose network gridlock over a single undetected routing error. That is the correct engineering behaviour, and it is the reason no aircraft were harmed. What made the disruption catastrophic was not the failure. It was that manual processing by human controllers could handle only a fraction of the traffic that the automated system had been carrying, and there was no way for the humans to make up the difference by working harder or faster. The system had been designed with four hours of buffered flight plans as a resilience margin. Four hours was not enough. Everything above the manual throughput ceiling had to be cancelled.
INSIGHT AND ANALYSIS
The phenomenon the NATS incident illustrates has been formally named in industrial engineering literature for over four decades. Lisanne Bainbridge published a paper called “Ironies of Automation” in the journal Automatica in 1983, arguing that as automated systems become more capable, human operators become progressively less able to intervene effectively when the automation fails.
The 2026 peer-reviewed medical literature calls this the automation paradox and points to aviation as the clearest domain in which it has been documented. The canonical case is Air France Flight 447 in 2009, where investigators concluded that the pilots who took manual control after the autopilot disconnected had accumulated too few recent manual flying hours to reason effectively about what the aircraft was doing.
The Nature Digital Medicine team has recently applied the same framework to hospital-based AI systems and found the same pattern — clinicians who use AI diagnostic tools for extended periods perform worse on the underlying diagnostic task when the AI is removed.
The automation paradox is the mechanism by which that principle breaks down in practice. If the human’s contribution to a normal-day decision is minimal because the automation handles it well, the human’s capacity to make the decision unassisted degrades over time. The organisation only discovers the degradation when the automation fails, at which point the failure is already public and consequential. The NATS incident is a live case study of exactly this trajectory. UK air traffic controllers are among the most rigorously trained operators in any industry. They still could not handle at manual pace what their automated system was handling on an ordinary Tuesday morning. That is not a criticism of the controllers. It is a description of what automation systematically does to any human process it substitutes for.
The question this raises for South African corporates is not theoretical. Every organisation of any size in this country has been running an automation programme for at least the past five years. Payment reconciliation. Customer service. Fraud detection. Credit decisioning. Call routing. Treasury operations. Compliance monitoring. Regulatory reporting. In each case, the business case for automation was written on efficiency gains and error reduction, which are real and defensible. In almost none of those cases was the business case written on manual reversion capacity, which is what actually matters when the automated system fails safely at 08:32 on a Tuesday morning.
The problem in South African financial services goes further than the general automation paradox. South African banks and insurers have moved core operations to cloud and SaaS platforms at pace over the past five years, and in doing so many have decommissioned the legacy manual screens and back-office processes that used to sit behind the automation as a manual fallback. If a core transaction scoring engine or authorisation platform fails safely tomorrow at your organisation, the question is not only whether the staff who used to work the manual process are still there and current. The question is whether the manual process still exists to work. In many cases, the answer is no. That is a deeper form of the same paradox, and it is now the state of a significant portion of the country’s financial services infrastructure. Most South African corporate risk registers do not carry manual reversion capacity as a line item. Most audit committee agendas have not asked about it. That is now, on the evidence of what happened in the UK last week, a materially incomplete governance posture.
IMPLICATIONS
For a South African board, the specific question that follows from the NATS incident is not “how do we protect our systems from rogue AI.” It is a different question, and it has three parts. First, what is our manual throughput ceiling — the number of transactions or decisions per hour our people can execute unassisted for each critical automated process. Second, how long can our organisation operate at that ceiling before the backlog produces systemic damage, whether that damage is liquidity failure, reputational default, or regulatory breach. Third, how many of the people in the organisation have ever actually run the manual process at business volume in live production, and are they still available to be called on when the automated system fails safely on the day it fails? Each of these questions has an honest answer inside every organisation. Very few boards currently know what the honest answer is, and in some cases the honest answer is that the manual process no longer exists at all.
The wider implication is that automation programmes need a category the current programme design does not typically include. Every large automation deployment should include a specific plan for how manual capability is maintained after the automation goes live. That plan will cost money, and it will feel like a duplication of what the automation was supposed to eliminate. It is not a duplication. It is the reason the organisation can survive the day the automation fails safely, and the price of not having it is what happened in the UK last week. In some domains, the plan will be a rotation of staff through manual work to keep their skills current. In others, it will be a documented manual-mode procedure with regular live drills. In some cases, it will be a contractual retainer with a service provider who can be called on for surge capacity in the specific manual mode. In every case, the plan should be visible in the risk register, should have named accountability, and should be tested at least annually. None of these is standard practice in South African corporate automation programmes today.
CLOSING TAKEAWAY
The NATS incident is not an argument against automation. Automation, including artificial intelligence, is going to keep growing across every industry in this country, and most of it is genuinely worth doing. What the NATS incident is an argument for is honesty about what automation removes from an organisation while it is delivering its benefits. It removes the human capacity to run the process manually, quietly and continuously, so that on the day the automated process fails safely — and safety-critical automation is engineered to fail safely, which means it will one day stop rather than pass ambiguous data through — the organisation has to make up the difference with the manual capacity it has kept. If the organisation has not kept any, or has kept only enough for a fraction of business volume, or has decommissioned the manual process entirely during a cloud migration, the failure becomes visible in the same way the NATS failure became visible. Publicly, expensively, and beyond the organisation’s ability to control. The board that walks into the next audit committee meeting with an AI strategy paper and no manual reversion capacity update is walking in with an incomplete answer to a question the UK air traffic control system has just asked on behalf of every organisation that automates its critical processes. Ask the question this quarter.
Johan Steyn is a prominent AI thought leader, speaker, and author with a deep understanding of artificial intelligence’s impact on business and society. He is passionate about ethical AI development and its role in shaping a better future. Find out more about Johan’s work at https://www.aiforbusiness.net




Comments