Frequency Is Not Risk
What line-observation data says about pilot error, and what it implies for the design of competency assessment
By Cedric Paillard, CEO, Amris Aviation
EXECUTIVE SUMMARY
Twenty years of Line Operations Safety Audit (LOSA) data describe what happens on ordinary revenue flights, and they contain a finding that most assessment systems are structurally unable to act on: the errors flight crews make most often are not the errors they handle worst.
Checklist errors are the single most common procedural error observed on the line, present on 26% of flights. They account for 5% of the errors crews mismanage. Manual handling and flight control errors account for 36%. A training priority list built from occurrence data and one built from consequence data are not the same list, and an assessment platform that records only whether a marker occurred cannot produce the second.
This paper sets out the evidence base in three layers — line observation, accident outcomes, and the specific behavioural gaps the record singles out — then maps the error families onto the ICAO competency framework, closing with what follows for assessment design. It also lists the widely circulated figures that the evidence does not support because several of them do not survive scrutiny by a customer safety department.
The central conclusion is narrow and practical. The competency framework is not the constraint: nine competencies and 73 observable behaviours already exist to trace a task failure to a root cause. The constraint is that most assessment records occurrence alone — not whether the error was trapped, and not what it turned into. Both numbers are present in the same observation. Only one is usually stored.
1. Two Populations, Two Questions
Accident data tells you how flights end. It is authoritative, comparable across years, and it is what almost every training priority in the industry is ultimately justified against. It also describes an extraordinarily rare population. IATA recorded 51 accidents across 38.7 million flights in 2025, a rate of 1.32 per million sectors.¹ Eight were fatal.
A competency framework, by contrast, has to grade the other 38.7 million. What happens on those flights is not visible in accident statistics by construction, and it is the LOSA that was designed to make it visible: trained observers riding the jumpseat on ordinary revenue flights, recording threats, errors and undesired aircraft states under an explicit no-jeopardy agreement. The archive now exceeds 20,000 observations at more than 70 airlines, accumulated since 1996.²
The two populations answer different questions, and the industry is much better at asking the first than the second. What ends a flight badly? — well understood. What happens on a normal Tuesday? — vaguer, and more consequential for anyone designing a marking scheme.
2. What Normal Operations Look Like
The LOSA picture is not a fleet of pilots making occasional mistakes. It is a routine working environment in which error is continuous and containment is the real variable.
Table 1 — Headline line-observation metrics
| Metric | Value | Basis |
| Flights with at least one crew error | ~80% | LOSA archive, 4,532 observations, 2002–2006³ |
| Mean errors per flight | ~3 | As above |
| Errors undetected or not responded to | 45% | As above |
| Errors mismanaged | 25% | Of which 6% led to a further error and 19% directly to an undesired aircraft state |
| Flights reaching an undesired aircraft state | ~33% | As above |
| Flights with at least one intentional non-compliance | 49% | Klinect, FSF IASS 2013, 20,000+ observations² |
| Flights carrying a checklist error | 26% | As above |
| ATC-related threat present | 61% of observations | As above |
Percentages are drawn from two separate archive reports and are not additive across rows.
Two of these deserve emphasis. The first is that 45% of observed errors are never detected or responded to at all. They are not mishandled; they are simply never noticed. The second is that of the quarter of errors that are mismanaged, roughly one in five converts directly into an undesired aircraft state.
Detection, in other words, carries most of the safety margin. An error that is caught is a non-event. The same error uncaught is the first link in a chain. And detection is among the least reliably assessed things in recurrent training: a review of nineteen US carrier simulator training plans found that only five specifically mentioned pilot monitoring.⁴
3. The Frequency–Consequence Inversion
Procedural errors dominate the count. Checklist errors are the most common procedural error observed on the line, followed closely by callout omissions and failures of SOP cross-verification; briefing errors are the least common of the family.³ Build a priority list from occurrence alone and checklist discipline sits at the top of it.
Filter the same dataset for the errors that were actually mismanaged, and the ordering inverts.
Table 2 — Composition of mismanaged errors
| Error type | Share | Note |
| Manual handling / flight control | 36% | Largest single share |
| Automation | 16% | |
| Systems, instruments, radio | 16% | |
| Checklist | 5% | Most frequent procedural error |
| Crew–ATC communication | 3% |
LOSA archive, 25 most recent audits, 4,532 observations, 2002–2006.³ The five named types do not exhaust the mismanaged population.
The mismanagement rates for individual event types make the same point with more force. These are conditional rates: given that the event occurred, how often did the crew fail to contain it?
Table 3 — Mismanagement rate by event type
| Event | Mismanaged |
| Speed deviation while hand flying | 81% |
| Automation error | 46% |
| FMS entry error | 35% |
| ATC transmission carrying three or more instructions | 19% |
| Pop-up aircraft malfunction | 16% |
| Challenging ATC speed clearance | 13% |
| Thunderstorm | 12% |
Klinect, FSF 66th International Air Safety Summit, 2013.²
The accident record corroborates the direction. IATA’s 2025 analysis identified aeroplane flight path management by manual control as a contributing factor in 29% of the year’s accidents, situation awareness and management of information in 29%, manual handling and primary flight control errors in 24%, and non-compliance with standard operating procedures in 20%.¹ Boeing’s current statistics place final approach and landing together at 48% of fatal accidents and 39% of onboard fatalities in 4% of flight time; landing alone accounts for 35% of fatal accidents in 1% of flight time.⁵
The practical implication is direct. A marker set weighted by observation frequency over-represents application of procedures. A set weighted by consequence over-represents manual flight path management. Neither weighting alone is a fair account of the operation, and the choice between them is currently made implicitly by what the assessment system happens to store.
4. Where Errors Occur, and When
Threats and errors do not cluster in the same phase of flight. Roughly 40% of threats arrive in predeparture and taxi-out, driven by airline-origin threats such as documentation, loading and dispatch. Roughly 40% of errors, and 55% of mismanaged errors, occur in descent, approach and landing.³
Table 4 — Phase distribution
| Phase | All threats | All errors | Mismanaged errors |
| Predeparture / taxi-out | ~40% | ~30% | Not stated |
| Descent / approach / landing | ~30% | ~40% | 55% |
| All other phases | ~30% | ~30% | Not stated |
Merritt & Klinect, 2006.³ Values are read from published charts and are stated as approximate.
The threat is planted on the ground and matures in the air. That has a direct consequence for scenario design in the scenario-based training phase: a line-oriented scenario that seeds its threats before pushback and lets them ripen into the approach reproduces the real distribution, while one that begins at the holding point does not. It also renders an entire error family invisible. Runway and taxi discipline is the one category in which pilot deviation is unambiguously the dominant cause — 61.9% of the 1,758 US runway incursions recorded in FY2024 were pilot deviations — and IATA reports that 48% of all accidents over the past decade relate to runway safety.⁶˒¹
5. Three Behaviours the Record Singles Out
5.1 Intentional Non-compliance Behaves as a Multiplier
Deliberate deviation from procedure was present on 49% of observed flights. Its significance is not the deviation itself but its correlation with everything else on the sheet: crews with no intentional non-compliance averaged 2.1 unintentional errors per flight, crews with one averaged 3.9, and crews with two or more averaged 7.5.²
The behaviours named in the archive are mundane rather than reckless: omitted altitude callouts, checklists performed from memory, failure to execute a mandatory missed approach, flight guidance changes made while hand flying, and taxi duties performed while the aircraft is still on the runway. Recorded as a session-level flag rather than only as a competency grade, intentional non-compliance is the single most predictive marker available to an observer.
5.2 Continuing an Unstable Approach is a Decision Failure
This is the largest single behavioural gap in the record, and the one where two credible sources give different-looking numbers for the same phenomenon. The FAA Air Carrier Training ARC, drawing on data from ten operators including six major Part 121 carriers, reports that roughly 10% of approaches exceed at least one stabilisation parameter at the 500 ft gate, and roughly 0.5% of those result in a go-around.⁷ The Flight Safety Foundation, using a tighter definition of unstable, found 3.5 to 4% of approaches unstable with 95 to 97% continued to landing.⁸
The figures are reconcilable: a single-parameter exceedance at a defined gate is a looser trigger than the FSF criterion, which mechanically enlarges the denominator and shrinks the compliance rate. Both should be cited with their thresholds; averaging them produces a number no one published.
The consequence side is unambiguous. The FSF study found that 83% of runway excursions over a sixteen-year period would have been avoided by a decision to go around, and that more than 80% of approach-and-landing accidents were similarly preventable. It also found that fewer than half of the pilots surveyed who had flown an unstable approach believed they would face any reprimand for landing off it.⁸
For assessment, the essential point is that the aircraft was flyable. The gate was called and the crew continued. Graded as a flight path deviation, the finding prescribes additional manual handling practice for a crew whose handling was not the problem. It belongs to problem solving and decision-making, with a share to the monitoring pilot who observed the deviation and did not escalate.
5.3 Monitoring is the Competency that Determines Everything Else
The 45% non-detection rate is the clearest statement of this in the line data. The accident record says the same: an NTSB review found monitoring errors present in 31 of 37 US airline accidents examined, and IATA’s 2025 analysis places situation awareness and management of information among the two most frequent contributing factors.⁴˒¹
A session in which every pilot-flying error was trapped and one in which none were trapped can produce identical flight paths, identical flight data traces, and identical narrative debriefs. They describe entirely different crews.
6. The Competency Framework as a Diagnostic Instrument
None of the above requires new theory. The ICAO competency framework, as published in PANS-TRG Amendment 7, comprises nine competencies and 73 observable behaviours, and exists precisely so that a task failure can be traced to a root cause rather than filed as an outcome.⁹ EASA’s examiner guidance states the principle directly: a lack of specific competencies may be identified as the root cause of the failure of the performance of a task. The manoeuvre is the symptom; the competency is the diagnosis.
Table 5 — The ICAO competency framework
| Abbr. | Official title | OBs | Errors that root here |
| KNO | Application of knowledge | 7 | Systems mis-operation |
| PRO | Application of procedures and compliance with regulations | 7 | Checklist, callout, non-compliance |
| COM | Communication | 10 | Briefing, readback, escalation |
| FPA | Aeroplane flight path management, automation | 6 | Mode selection, FMS entry |
| FPM | Aeroplane flight path management, manual control | 7 | Handling, landing phase |
| LTW | Leadership and teamwork | 11 | Failure to challenge |
| PSD | Problem solving and decision-making | 9 | Unstable approach continuation |
| SAW | Situation awareness and management of information | 7 | Monitoring, energy state |
| WLM | Workload management | 9 | Interruption-driven breaks |
ICAO Doc 9868 (PANS-TRG), 3rd edition 2020, Amendment 7. Titles as reproduced in IATA guidance material and cross-checked against the EASA Flight Examiner Manual.⁹˒¹⁰
IATA constructs the competency grade from three dimensions: how many of the observable behaviours were demonstrated when required, how often they were demonstrated, and what the outcome of threat and error management was.¹⁰ The third dimension is the one that matters here, because unlike the first two it has an external referent. The LOSA archive already quantifies what proportion of errors are trapped, mismanaged, and converted into an undesired aircraft state. A grading boundary whose definition is “a reduction in safety margin” is describing something the line data measures directly.
The instructor workflow that IATA prescribes — observe performance, record details, classify observations against the observable behaviours, and assess by determining root causes according to the competency framework — already contains the step at which this evidence would be applied. Classification is where an error family becomes a competency, and it is the step at which the distinction between a compliance failure and a workload failure that looks like one is either captured or lost.
That distinction is not academic. A checklist run from memory in a quiet cockpit is an application-of-procedures finding. The same checklist abandoned after an unexpected ATC call is a workload management finding presenting as a procedural one. Two identical entries on a marking sheet; two entirely different remedies.
7. What Follows for Assessment Design
Five consequences follow from the evidence above. None of them require a change to the regulatory framework, and none are technically demanding.
- Record two rates per marker, not one. Occurrence and mismanagement diverge sharply enough that a system reporting only frequency will direct an operator’s training budget at checklist discipline. Both values are present in the same observation; conventionally only the first is stored.
- Capture the error-to-undesired-state transition explicitly. It is the third dimension of the IATA grade and the only one with an external referent. Recording whether an error was trapped, mismanaged, or converted into an undesired state gives the boundary between grades a defensible anchor rather than an instructor’s judgement alone.
- Grade the monitoring pilot on a separate sheet. The non-detection rate is the single largest lever in the line data, and it is invisible in a marking scheme that grades the flight path rather than the crew.
- Treat marking distribution as a calibration signal. If an instructor’s findings cluster on manual handling while the population data indicates that roughly half of line errors are procedural, that divergence is a calibration input available without placing a second observer in the room — and it survives audit in a way that subjective moderation does not.
- Weight scenario design by phase. Threats belong on the ground and errors in the descent. A scenario that inverts this trains against a distribution that does not occur, and a scenario that begins at the holding point removes an entire error family from assessment.
One caveat applies to portability. The mapping above is expressed in ICAO terms and transfers directly to EASA states, where evidence-based training is implemented through ORO.FC.231. The FAA’s Advanced Qualification Program is proficiency-based, data-driven and scenario-centred, and is structurally comparable, but it does not adopt the nine-competency framework, the competency abbreviations, or the five-point competency scale; AQP operators define their own terminal and enabling objectives.¹¹ Any competency-referenced assessment intended for a US Part 121 operator requires a translation layer to that operator’s own qualification standards.
8. Claims the Evidence Does Not Support
Several widely circulated figures in this area do not survive checking. They are listed here because each is the kind of claim a customer safety department will test.
- “X% of accidents are caused by pilot error.” IATA publishes no such figure. Its contributing-factor percentages are reported individually and explicitly exceed 100%, because accidents carry multiple factors. They cannot be summed into an aggregate.
- A percentage split of the original five LOSA error categories. The five-category taxonomy — intentional non-compliance, procedural, communication, proficiency, operational decision — was superseded by a three-category scheme before large-archive reporting matured. No archive-wide split of the five categories appears to have been published.
- Loss of control in flight as the current leading killer. There were no loss-of-control-in-flight accidents in 2025, the second consecutive year. LOC-I remains a legitimate training priority and the historic leading fatality category; presented as a current statistic it is incorrect and checkable.
- “47% of fatal general aviation accidents involve loss of control.” The NTSB figure covers 2008–2014 and was published in 2016. The Most Wanted List was retired in December 2023 and no direct successor to that statistic exists.
- An averaged go-around compliance rate. The 3% and 0.5% figures use materially different definitions of unstable. Both should be cited with their thresholds.
- Citing ICAO Doc 9683 for threat and error management, or Doc 10111 for competencies. Doc 9683, the Human Factors Training Manual, builds on the SHEL model and does not contain the LOSA taxonomy. Doc 10111 is the manual on cabin electronic flight bags. The competency framework is Doc 9868 (PANS-TRG); the evidence-based training manual is Doc 9995.
- Treating the Doc 9803 figures as archive-wide. Those values — 85% of crews making at least one error, errors in 74% of segments, two errors per segment — come from ICAO Doc 9803 and describe a single airline observed between 1996 and 1998. The archive-wide equivalents are approximately 80% of flights and three errors per flight.
9. Conclusion
The training industry has spent two decades building a competency vocabulary precise enough to describe why a crew failed rather than merely that they did. The vocabulary is complete. Nine competencies, 73 observable behaviours, a grading model with a threat-and-error dimension, and a regulatory route in EASA states.
What has not kept pace is the record. Assessment systems overwhelmingly capture whether a marker occurred. The line data has been saying for twenty years that occurrence is the less informative of the two available numbers — that what distinguishes a safe operation from an unsafe one is not how many errors were made but how many were caught, and what the uncaught ones became.
That second number does not require new instrumentation, new regulation, or new observation. It requires only that the observer record it and the system keep it.
References
1 IATA, Annual Safety Report 2025 — Executive Summary and Safety Overview, published March 2026. https://www.iata.org/contentassets/a8e49941e8824a058fee3f5ae0c005d9/safety-report-executive-summary-and-safety-overview_2025.pdf
2 Klinect, J., presentation to the Flight Safety Foundation 66th International Air Safety Summit, Washington DC, 29–31 October 2013; reported as “Intentionally Noncompliant”, AeroSafety World. Basis: more than 20,000 LOSA observations at more than 70 airlines since 1996. https://flightsafety.org/asw-article/intentionally-noncompliant/
3 Merritt, A. & Klinect, J., Defensive Flying for Pilots: An Introduction to Threat and Error Management. University of Texas Human Factors Research Project / The LOSA Collaborative, 12 December 2006. Basis: 4,532 observations from the 25 most recent LOSAs, 2002–2006. https://skybrary.aero/sites/default/files/bookshelf/1982.pdf
4 US Department of Transportation Office of Inspector General, Report AV-2016-013, Enhanced FAA Oversight Could Reduce Hazards Associated With Increased Use of Flight Deck Automation, 7 January 2016. Also the source for the NTSB finding on monitoring errors. https://www.oig.dot.gov/sites/default/files/FAA%20Flight%20Decek%20Automation_Final%20Report%5E1-7-16.pdf
5 Boeing, Statistical Summary of Commercial Jet Airplane Accidents, Worldwide Operations 1959–2025, 57th edition, April 2026. Phase-of-flight data covers 2016–2025. https://www.boeing.com/content/dam/boeing/v2/safety/statsum.pdf
6 US Department of Transportation Office of Inspector General, Report AV2025026, FAA Has Taken Steps To Prevent and Mitigate Runway Incursions, 12 March 2025. https://www.oig.dot.gov/sites/default/files/library-items/FAA%20Runway%20Incursions%20Final%20Report_3.12.25.pdf
7 FAA Air Carrier Training Aviation Rulemaking Committee, Recommendation 24-2: Stabilized Approach Policy, 19 September 2024. https://www.faa.gov/about/office_org/headquarters_offices/avs/offices/afx/afs/afs200/afs280/act_arc/act_arc_reco/ACT_ARC_Recommendation_24-2.pdf
8 Blajev, T. & Curtis, W., Go-Around Decision-Making and Execution Project: Final Report to Flight Safety Foundation, 2017. Survey base 2,340 pilots and 128 managers. https://flightsafety.org/wp-content/uploads/2017/03/Go-around-study_final.pdf
9 ICAO, Doc 9868, Procedures for Air Navigation Services — Training (PANS-TRG), 3rd edition 2020, incorporating Amendment 7, applicable 5 November 2020. Appendix 1 to Chapter 1 carries the competency framework.
10 IATA, Competency Assessment and Evaluation for Pilots, Instructors and Evaluators — Guidance Material, 4th edition, 2025; and Evidence-Based Training Implementation Guide, edition 2, effective January 2024. https://www.iata.org/contentassets/c0f61fc821dc4f62bb6441d7abedb076/competency-assessment-and-evaluation-for-pilots-instructors-and-evaluators-gm.pdf
11 FAA, Advisory Circular 120-54A, Advanced Qualification Program, 23 June 2006, Change 1 (2022); 14 CFR Part 121 Subpart Y. EASA implementation: ORO.FC.231, introduced into Regulation (EU) 965/2012, with AMC/GM at ED Decision 2021/002/R.
NOTE ON METHOD
Every figure in this paper is attributed to a named source with a publication year. Where a value was read from a published chart rather than a numeric table — the phase distribution in Table 4 and the approximate three-way split of errors by type — it is stated as approximate. The competency titles, descriptions and observable behaviour counts are taken from IATA’s 2025 guidance material, which states that it reproduces the ICAO Doc 9868 Amendment 7 framework, and were cross-checked against the EASA Flight Examiner Manual; the two agree word for word. ICAO’s own text was not read directly. Figures that could not be verified against a primary or reputable secondary source have been excluded rather than estimated, and the principal exclusions are listed in section 8.