Governance-First AI in Aviation Training
What the Enterprise AI ROI Crisis Teaches Us About Building Trustworthy Training Intelligence
The ROI Reality Check That Every Industry Needs to Hear
In a recent analysis that has resonated across the technology sector, AI strategist Aliya Nur Babul laid bare an uncomfortable truth about enterprise AI adoption. Her thesis is simple and devastating: the gap between what AI vendors promise and what organisations actually experience after 90 days in production is not a minor discrepancy. It is a structural failure in how enterprises approach artificial intelligence.
Her research reveals that after just three months in production, enterprise AI budgets typically break down into four largely unplanned cost centres:
- API fees consuming roughly 30% of spend (driven by error loops and hallucinated tool chains),
- Human supervision absorbing 25% (operators manually validating outputs and chasing evidence trails),
- Rework accounting for 20% (fixing what autonomous agents break through untraceable decisions), and
- Governance overhead consuming the remaining 25% (audit preparation, compliance mapping, and drift monitoring).
The conclusion is stark: what boards were promised as a productivity multiplier has, in many cases, become a budget black hole. And the organisations that succeed are not the ones chasing the flashiest demos. They are the ones that build governance-first, treating AI as accountable infrastructure with measurable cost controls rather than experimental side projects.
The winners build governance-first, treating AI as accountable infrastructure with measurable cost controls — not experimental side projects. — Aliya Nur Babul, AI Operator & Strategist
For those of us in aviation, this analysis is not merely relevant. It is prophetic. Because the enterprise AI pitfalls that Babul describes — the supervision tax, the governance reckoning, the rework spiral — are precisely the traps that aviation training technology must avoid. The stakes in our industry are not quarterly earnings. They are human lives.
The Aviation Training Parallel: Same Pitfalls, Higher Stakes
The Supervision Tax in the Cockpit
Babul introduces a concept she calls the Supervision Tax: the hidden cost of humans who must continuously validate, correct, and intervene in AI-driven workflows. She describes it as “automation theatre” — the illusion of efficiency masking the reality that human operators are spending as much time supervising the AI as they would have spent doing the work themselves.
Aviation training departments will recognise this pattern immediately. For decades, the industry has invested in digital tools that promise to streamline assessment and reporting, only to discover that instructors end up doing double work: completing the digital workflow and then separately documenting their actual observations in notes, emails, or verbal debriefs because the system does not capture what they really need to record.
The root cause is the same in both cases. When AI systems are designed to replace human judgment rather than support it, the humans do not disappear from the process. They simply shift from doing productive work to supervising a system that does not fully understand their domain. The supervision tax is not a technology problem. It is an architecture problem.
The Rework Spiral in Training Data
Babul identifies rework — at 20% of enterprise AI budgets — as the cost of fixing what autonomous agents break through untraceable decisions. In enterprise software, this manifests as corrupted data, incorrect reports, or automated actions that require manual reversal.
In aviation training, the equivalent is the subjective assessment problem. When instructor observations remain as unstructured opinions rather than classified, evidence-bound data, the downstream consequences are severe. A pilot’s competency profile becomes a collection of disconnected snapshots that cannot be trended, compared, or audited. When a training decision is challenged — by a regulator, by a pilot, or by an incident investigator — the data trail is incomplete or inconsistent. The rework is not just costly. In aviation, it can be catastrophic.
The Governance Reckoning at 30,000 Feet
Perhaps Babul’s most striking insight is what she calls the Governance Reckoning: the discovery that retrofitting compliance and auditability into AI systems that were not designed for it costs two to three times the original build. She uses the iceberg metaphor — 90% accuracy above the waterline, but beneath the surface lie full manual audits, compliance violations, and customer churn.
Aviation operates in one of the most heavily regulated environments on earth. ICAO competency frameworks, EASA oversight requirements, and national authority audit expectations create a governance landscape that makes the EU AI Act look straightforward by comparison. Any AI system deployed in pilot training that was not designed with governance at its core will inevitably face Babul’s reckoning — and in aviation, the consequences of governance failure are measured not in budget overruns but in safety margins.
Amris AI: Built Governance-First from Day One
It is against this backdrop that the design philosophy behind the Amris Competency Management Platform (originally Amelia AI), developed by The Airline Pilot Club (APC) – now Amris Aviation – becomes particularly significant. Amris CMP was not conceived as an AI tool that would be retrofitted with governance later. It was architected from its foundation as a governance-first, human-in-the-loop intelligence layer for pilot competency management.
Every design decision in Amris CMP directly addresses one of the enterprise AI failure modes that Babul identifies. Understanding these parallels reveals why purpose-built, domain-specific AI consistently outperforms general-purpose tools in regulated environments.
Eliminating the Supervision Tax Through Purposeful Human-in-the-Loop Design
Where enterprise AI tools create a supervision tax by requiring humans to validate AI outputs they did not generate, Amris CMP inverts the model entirely. The platform’s ORCA workflow — Observe, Record, Classify, Assess — is designed so that the human instructor remains the source of truth at every stage. The instructor observes pilot behaviour, records their observations, classifies them against ICAO-aligned competency baselines, and makes the assessment.
Amris’s role is not to replace any of these steps. It is to structure, validate, and persist them. When an instructor captures a behavioural observation, Amris CMP ensures it maps correctly to the relevant competency framework. When patterns emerge across multiple sessions, Amris surfaces them. When grading consistency needs to be calibrated across a department, Amris provides the data to make that calibration objective.
This is the fundamental difference between automation theatre and genuine augmentation. The instructor’s time is not consumed by supervising an algorithm. It is enhanced by a system that gives their expert judgment structure, context, and permanence. The result is a measurable reduction in administrative burden — approximately 20-25% less paperwork time — without the hidden supervision costs that plague general-purpose AI deployments.
Eliminating Rework Through Traceable, Evidence-Bound Decisions
Babul’s rework problem stems from AI systems making untraceable decisions that humans must then diagnose and reverse. Amris CMP eliminates this failure mode by design: the platform does not make autonomous decisions about pilot competency. Every assessment is instructor-originated, digitally captured, and mapped to a defined framework.
This creates what Babul would recognise as a fully traceable spend model applied to training data. Every observation has provenance: who recorded it, when, against which competency indicator, and within what training context. When a decision is queried — whether by a line manager, a regulator, or a review board — the evidence trail is complete and defensible. There is no rework because there are no black-box decisions to reverse.
The competency drift detection capability illustrates this further. Rather than an AI autonomously flagging pilots as deficient (which would create exactly the rework and contestation problems Babul describes), Amris CMP surfaces longitudinal patterns that instructors can interpret in context. The system provides the signal. The human provides the judgment. The result is defensible, not disputable.
Governance as Architecture, Not Afterthought
Babul’s governance reckoning — the 2-3x cost multiplier of retrofitting compliance — is perhaps where Amris’s design philosophy diverges most dramatically from the enterprise AI failures she documents. Amris CMP was built from inception to satisfy the governance requirements of one of the world’s most demanding regulatory environments.
Compliance is not a feature that was added to Amris CMP. It is the architecture itself. The ORCA workflow is designed to produce CBTA/EBT-ready evidence. The data model maps directly to ICAO competency frameworks. The audit trail is continuous and complete from the moment a pilot enters the Amris Aviation ecosystem through recruitment, through academy training, and into the Airline Ready Pilot Pool (ARPP).
The question for airline leadership is no longer whether to adopt AI in training. It is whether the AI they adopt was built to withstand the scrutiny that aviation demands.
When Babul argues that forward-thinking operators should demand a 90-day cost waterfall, traceable spend, auditable controls, and measurable ROI, she is describing exactly the accountability model that Amris provides to airline training departments. Every data point has an origin. Every assessment has an evidence basis. Every competency trajectory can be audited from initial selection through to line operations.
The Due Diligence Shift: From Demo-First to Governance-First
Babul concludes her analysis with what she calls the Due Diligence Shift: the transition from evaluating AI tools based on impressive demos to evaluating them based on what happens after 90 days in production. The old question was simple and seductive: what is your demo ROI? The new question is harder but essential: show me your cost waterfall after 90 days.
For aviation, this shift has even deeper implications. Airlines and training organisations evaluating AI-enabled competency management tools should be asking not just about cost waterfalls but about evidence integrity, regulatory defensibility, and long-term data continuity. A tool that delivers an impressive demo but produces unstructured, unauditable data in production is not merely expensive. In aviation, it is a liability. The Amris CMP platform and the broader Amris Aviation ecosystem — spanning evidence-based recruitment, AI-powered training intelligence, structured academy oversight, and the curated ARPP talent pipeline — represent what a governance-first approach looks like when applied with domain expertise to one of the world’s most safety-critical industries.
A Lesson That Crosses Industries
Babul’s analysis serves as a powerful reminder that the value of AI is not determined by what it can do in a controlled demonstration. It is determined by what it costs, what it breaks, and what it can prove after sustained deployment in complex, real-world environments.
Aviation has always understood this instinctively. Every system that enters a cockpit is tested not for its best-case performance but for its failure modes. The same standard should apply to any AI system that touches pilot competency, training assessment, or talent management.
The organisations that will lead aviation’s next chapter are those that choose AI partners who understand this. Not the vendors with the most impressive demos, but the ones who can show you, transparently, what their system looks like after 90 days, 900 days, and 9,000 assessments in production. Governance-first is not a constraint. It is the foundation of trust.
For more information on how Amris Aviation is pioneering governance-first AI in aviation training, visit https://amrisaviation.com