The Cognitive Limits of CBTA
Why Working-Memory Capacity and Statistical Power Make AI-Assisted Observation a Regulatory Requirement, Not an Enhancement
A quantitative analysis of instructor cognitive load and grading reliability in the ICAO/EASA Competency-Based Training and Assessment framework
Abstract
Competency-Based Training and Assessment (CBTA) — as codified by ICAO Doc 9995, EASA AMC1 ORO.FC.231 and the IATA EBT framework — specifies an assessment architecture of nine pilot competencies, seventy-three Observable Behaviours (OBs), and a five-point grading scale calibrated against frequency claims (“rarely”, “occasionally”, “regularly”, “very often”, “almost always”). This paper applies two well-established quantitative tools — Cowan’s working-memory capacity model and Wilson’s binomial confidence interval — to ask whether human instructors, observing in real time, can deliver the assessment fidelity that the framework specifies.
The findings are unambiguous. During the four highest-criticality phases of flight (Take-off, Climb, Approach, Landing), instructor cognitive load reaches 4.8× the established working-memory capacity limit. A single four-hour simulator session delivers approximately 3.5 demonstrations per OB, against a statistical requirement of 25–97 demonstrations per OB for defensible grading. Per-OB CBTA grading from a single session is, in formal statistical terms, indefensible. Per-competency grading is reliable only at coarse precision tiers and only after multiple sessions are pooled.
These limits are not deficiencies of instructor skill. They are arithmetic. The implication is that the CBTA framework, as specified, is not deliverable by unaided human observation — it requires AI-assisted continuous observation and cross-session evidence pooling to meet the statistical and cognitive prerequisites the framework itself implies. We argue that this reframes AI-assisted assessment from an efficiency tool to a regulatory necessity.
1. Introduction
1.1 The promise of CBTA
The transition from task- and hours-based pilot training to competency-based training and assessment represents one of the most significant pedagogical shifts in commercial aviation since the introduction of the simulator. Where conventional checking verifies whether a pilot performed manoeuvres within published tolerances, CBTA examines how and why performance occurred — assessing the underlying competencies (knowledge application, situation awareness, leadership and teamwork, workload management, and others) that produce reliable behaviour across novel operational contexts.
The shift is supported by ICAO Doc 9995, EASA AMC1 ORO.FC.231, and IATA’s Evidence-Based Training (EBT) implementation guide. At its core, CBTA assesses pilots against nine technical and non-technical competencies, each manifested through a defined set of Observable Behaviours (OBs). The full framework specifies seventy-three OBs spanning the entire flight profile, graded on a five-point scale where Grade 3 (“adequate”) is the IATA-recommended minimum acceptable performance.
1.2 The problem this paper addresses
CBTA’s assessment architecture is theoretically elegant and operationally formidable. The framework specifies that an instructor, observing a pilot in real time, must classify the frequency and quality of seventy-three discrete behaviours, integrated against the threat-and-error environment, while simultaneously running the assessment grammar (How Many × How Often × Outcome of TEM) for each of nine competencies.
This paper asks a question the published CBTA literature has not directly confronted: is this cognitively and statistically possible?
The question is not whether instructors are skilled enough. The most experienced Type Rating Examiner in the industry is bounded by the same human cognitive architecture as a first-year instructor. The question is whether the framework’s specified observation matrix exceeds the working-memory capacity of the human assessor, and whether a typical simulator session generates enough OB demonstrations for the resulting grades to clear basic statistical reliability thresholds.
1.3 Methodology and scope
We apply two independently established quantitative models to the CBTA framework as documented:
- Cowan’s working-memory capacity model (Cowan 2001, 2010), which establishes that the focus-of-attention capacity in normal adults averages approximately four chunks of information.
- Wilson’s binomial confidence interval (Wilson 1927; Brown, Cai & DasGupta 2001), which gives the sample size required to estimate a proportion with a target precision.
Section 2 establishes the cognitive science. Section 3 estimates the chunk load required by CBTA assessment for each phase of flight. Section 4 derives the statistical sample size required for grades to be meaningful. Section 5 reconciles these requirements against the demonstrations a typical simulator session actually delivers. Section 6 sets out the implications for assessment validity, regulatory defensibility, and the role of AI-assisted observation. Section 7 concludes.
Throughout, we use the term “instructor” to refer to any qualified CBTA assessor — Type Rating Instructor, Type Rating Examiner, EBT Instructor or Line Check Pilot. The arguments apply uniformly to all assessment roles.
2. Working Memory and the Chunk
2.1 From Miller to Cowan
In 1956, George A. Miller published “The Magical Number Seven, Plus or Minus Two,” arguing that human short-term memory could hold approximately seven discrete items. The figure entered psychology textbooks and, from there, the practical assumptions of instructional design across many fields, including aviation training.
Miller’s estimate is now considered an over-estimate. Beginning with his 2001 paper “The Magical Number 4 in Short-Term Memory,” and refined in subsequent work (Cowan 2005, 2010; Gilchrist, Cowan & Naveh-Benjamin 2008), Nelson Cowan demonstrated that when experimental conditions prevent participants from rehearsing or chunking, working memory capacity is closer to four items. Cowan’s review covered verbal and non-verbal materials, visual and auditory presentation, single and dual-task conditions, and attended versus unattended stimuli; the four-item limit appeared consistently across them. Cowan’s 2001 paper has been cited over 3,500 times and provides the formula now most commonly used to estimate capacity from array recognition tasks.
Subsequent work (Rouder et al. 2008; Shipstead, Lindsey, Marshall & Engle 2014; Adam, Mance, Fukuda & Vogel 2015) has corroborated the four-chunk limit, with individual variation typically ranging from three to five chunks. Cowan’s limit is now the consensus working assumption in cognitive psychology.
2.2 What a chunk is
A chunk is a meaningfully integrated unit of information held as a single entity in the focus of attention. Crucially, the size of a chunk depends on the expertise of the person holding it. A novice in physics holds the seven concepts of vector addition, vector subtraction, displacement, velocity, speed, acceleration and force as seven separate items; an expert holds them as one chunk because they have been compiled into a single integrated representation in long-term memory and can be activated together (Chase & Simon 1973; Ericsson & Kintsch 1995).
This is why expertise increases effective working-memory capacity without violating Cowan’s limit. The expert still holds four chunks; each chunk just contains more information. Chunking is the mechanism by which experience translates into superior performance under cognitive load.
2.3 The cognitive-load implication
From Cowan and from chunking theory follows the central tenet of Sweller’s Cognitive Load Theory (Sweller 1988; Sweller, van Merriënboer & Paas 1998, 2019): instructional design must respect working-memory capacity. Tasks that exceed the capacity of the focus of attention are not performed; they are sampled, with whatever is outside the focus of attention at that moment lost or reconstructed afterwards from cues in long-term memory. The reconstruction is biased by recency, salience, and the assessor’s prior expectations — known sources of error in any human observational protocol.
If a task requires the simultaneous tracking of more than four chunks, performance does not degrade gracefully. It fragments. The observer attends to a subset of the available signal and reconstructs the rest. The observer is rarely aware that this is what they are doing.
3. The Chunk Load of CBTA Assessment
3.1 What the instructor must hold concurrently
To execute a CBTA assessment in real time, the instructor must maintain in the focus of attention, at every moment of the active phase, the following streams of information:
- 1. Active sub-task identity — what the pilot is supposed to be doing right now. The framework specifies between four and ten tasks per phase, decomposed into fifteen to forty-two sub-tasks.
- 2. Observable Behaviours in scope — the OBs across the nine competencies that are relevant to the active sub-task. Approximately seven OBs map to a typical sub-task.
- 3. Threat and Error Management state — active threats, errors committed, and any undesired aircraft state (UAS). In low-workload phases this involves two or three items; in high-workload phases, six to eight.
- 4. Crew streams — what each crew member (Pilot Flying, Pilot Monitoring, and where applicable Cabin Crew or ATC) is doing, since multi-crew operation specifies separate OB Assessment Guides.
- 5. Running competency tally — the How Many × How Often × Outcome of TEM evaluation for each of nine competencies, the lowest of which determines the grade.
- 6. Word-picture grade anchors — the linguistic descriptors that distinguish Grade 3 (“adequate”) from Grade 4 (“effective”) for each of nine competencies. We assume these are compiled into long-term memory for an experienced assessor and excluded from the working-memory budget; this is a conservative assumption favourable to the unaided-instructor case.
3.2 Three levels of chunking
The same raw item count produces different working-memory loads depending on how effectively the assessor chunks. We model three levels:
- Novice: each item is its own chunk; no compilation has occurred.
- Typical expert: OBs are chunked by competency cluster, TEM threats are paired into threat-error pairs, crew members are tracked as separate streams.
- Optimal expert: hierarchical super-chunks (“phase context”, “TEM mental model”, “competency scoreboard”, “exception watch”) compress everything into four to six high-level units.
3.3 Per-phase chunk load
Applying this model to the CBTA framework, with task and sub-task counts taken directly from the published framework structure, yields the following load estimates:
| Phase | Tasks | Sub-tasks | Concurrency | TEM density | Typical-expert chunks | Overload vs Cowan limit (4) |
| Pre-flight | 7 | 34 | Sequential | Low | 14 | 3.5× |
| Take-off | 8 | 32 | Parallel | High | 19 | 4.8× |
| Climb | 7 | 27 | Parallel | Medium | 17 | 4.2× |
| Cruise | 7 | 27 | Sequential | Low | 14 | 3.5× |
| Descent | 8 | 32 | Parallel | Medium | 17 | 4.2× |
| Approach | 10 | 42 | Parallel | High | 19 | 4.8× |
| Landing | 4 | 15 | Parallel | High | 19 | 4.8× |
| Post-flight | 6 | 21 | Sequential | Low | 14 | 3.5× |
Table 1. Per-phase chunk load for CBTA assessment, computed against Cowan’s working-memory capacity limit of four chunks. Highlighted phases (Take-off, Approach, Landing) are the highest-criticality phases in commercial flight, where industry incident statistics concentrate.
3.4 What this means
In the lightest phases (Cruise, Pre-flight, Post-flight), instructor cognitive load is approximately 3.5× the established working-memory capacity. In the highest-criticality phases (Take-off, Approach, Landing), the figure reaches 4.8×. These are the phases where assessment matters most because they are where commercial-aviation incidents disproportionately occur, and they are precisely the phases where the assessor’s working memory is most overloaded.
This is not a marginal overload. At 4.8× capacity, the assessor is not making fine-grained discrimination between Grade 3 and Grade 4 demonstrations of, say, the seventh OB under Workload Management. The assessor is sampling a subset of the available signal and reconstructing the rest after the fact. This is not an indictment of instructor skill; it is the predictable consequence of asking a four-chunk system to track nineteen chunks of state.
4. The Statistical Reliability of CBTA Grades
4.1 The grade is a frequency claim
The CBTA grading scale, as set out in the framework’s Assessment Process, is structured around three dimensions: How Many (of the relevant OBs were demonstrated), How Often (the demonstrated OBs occurred during the assessable period), and Outcome of TEM (whether the result was unsafe, safe, or enhancing). The grade is the lowest of the three. Two of the three dimensions are explicit frequency claims, and the framework provides linguistic anchors for each:
| Grade | Performance label | How Many | How Often | Implied proportion p |
| 1 | Ineffectively | few / hardly any | rarely | p ≤ 20% |
| 2 | Minimal acceptable | some | occasionally | 20% < p ≤ 50% |
| 3 | Adequately | many | regularly | 50% < p ≤ 75% |
| 4 | Effectively | most | very often | 75% < p ≤ 90% |
| 5 | Exemplary manner | all / almost all | always / almost always | p > 90% |
Table 2. CBTA grade scale translated to demonstration proportions. Grade 3 (“Adequately”, highlighted) is the IATA-recommended target for end-of-EBT-module performance.
4.2 The sample size required
Once the grade is recognised as a proportion estimate, the question of reliability becomes a textbook problem. To estimate a proportion p with 95% confidence at a given precision (the half-width of the confidence interval), the required sample size n is approximated by the Wilson formula:
n ≈ (1.96² × p × (1 − p)) / HW²
Worst-case sample sizes (at p = 0.5, where variance is maximised) for each precision tier are as follows:
| Reliability tier | CI half-width | Use case | n per OB | n across all 73 OBs |
| Strict | ±5pp | Distinguishes adjacent grades near a band boundary; research-grade | 385 | 28,105 |
| Standard | ±10pp | ICAO-defensible; type-rating final assessment, line check | 97 | 7,081 |
| Pragmatic | ±15pp | Operational minimum; recurrent training, formative feedback | 43 | 3,139 |
| Indicative | ±20pp | Coarse tier classification only; not defensible | 25 | 1,825 |
Table 3. Sample size required per OB and across the 73-OB framework, by reliability tier. Standard tier is the minimum that supports defensible licence-impact decisions.
4.3 What a session actually delivers
A typical four-hour simulator session, with approximately 70% of duration available for active assessment (excluding briefings, repositions, and freezes), provides 168 active minutes. At an estimated rate of 1.5 OB-demonstration opportunities per active minute — a deliberately conservative figure pending calibration against operational ORCA pipeline data — a session generates approximately 252 demonstrations in total.
Distributed across the 73-OB framework, this is approximately 3.5 demonstrations per OB per session. Distributed across the 9 competencies (where evidence pools across the OBs of each competency), this is approximately 28 demonstrations per competency per session.
4.4 Sessions required for defensible grading
| Reliability tier | CI half-width | Sessions for per-OB grading | Sessions for per-competency grading |
| Strict | ±5pp | 112 | 13.8 |
| Standard | ±10pp | 28 | 3.5 |
| Pragmatic | ±15pp | 12.5 | 1.5 |
| Indicative | ±20pp | 7.2 | 0.9 |
Table 4. Sessions required for grades to reach each reliability tier, given a typical four-hour simulator session yielding ~3.5 demonstrations per OB and ~28 per competency.
4.5 The arithmetic conclusion
Three findings fall directly out of the table:
- Per-OB grades from a single sim session are statistical noise. A single session delivers ~3.5 demonstrations per OB. Even the loosest defensible precision tier requires 25 demonstrations per OB. Reaching it takes seven sessions; reaching ICAO-defensible Standard tier takes twenty-eight sessions per OB.
- Per-competency grading across a Type Rating is defensible. A 32-session type rating accumulates approximately 900 demonstrations per competency, comfortably past the Standard tier and approaching Strict tier. This is consistent with what regulators currently audit: competency-level grades, not per-OB grades.
- EBT recurrent training (typically two sessions per six-month cycle) reaches Pragmatic-tier per-competency reliability per cycle. It does not reach defensible per-OB reliability under any realistic scheduling, which is why the framework wisely targets Grade 3 (“adequate”) and treats finer discrimination as a longitudinal rather than a per-session product.
The CBTA framework’s per-OB granularity is theoretically beautiful and operationally undeliverable by unaided human observation. The framework specifies a level of measurement precision that the available human bandwidth — both cognitive and statistical — cannot supply.
5. The Two Limits Compound
Sections 3 and 4 establish two independent constraints. They are not independent in their effect on the assessor.
5.1 The cognitive bottleneck reduces the effective sample size
The statistical analysis of Section 4 assumes the assessor observes every OB demonstration that occurs. This is an upper bound on what the assessor sees. The cognitive analysis of Section 3 establishes that during the highest-criticality phases the assessor is operating at 4.8× working-memory capacity, which means the assessor is necessarily sampling. The effective number of OB demonstrations actually observed and registered is strictly less than the number that occur — perhaps half, perhaps less.
If the 3.5 demonstrations per OB per session figure overstates the registered sample by even a factor of two, then the sessions required to reach each reliability tier double. A single session no longer reaches Indicative tier per-OB; it reaches nothing per-OB. Per-competency reliability also halves.
5.2 The cognitive bottleneck biases the sample
Worse, the sampling is not random. The assessor’s attention is drawn to salient, recent, or expectation-confirming events, which means the demonstrations that enter the sample are systematically different from those that do not. This is not a ceiling-effect problem; it is a measurement-validity problem. The grade derived from such a sample reflects the assessor’s attention pattern as much as the pilot’s performance.
5.3 The cognitive bottleneck cannot be aggregated across sessions
Pooling across sessions is the natural statistical remedy: if a single session delivers 3.5 demonstrations per OB, four sessions deliver 14. But pooling requires the assessor to remember demonstration counts across sessions — a multi-week working-memory task that exceeds Cowan’s limit by orders of magnitude. In practice, demonstrations from prior sessions are reconstructed from training records, narrative debrief notes, and the assessor’s general impression. They are not aggregated as a quantitative sample.
The two limits are multiplicative. Working-memory capacity caps what can be observed in real time. Statistical power caps how that observation translates into a defensible grade. A solution to either limit alone is insufficient; a solution must address both.
6. AI-Assisted Observation as a Regulatory Necessity
6.1 What changes when continuous observation is mechanised
The cognitive bottleneck is a property of the human assessor, not of the observation task. A system that observes every OB on every demonstration without sampling fatigue, that registers each demonstration as a structured data point, and that aggregates across sessions automatically does not face the same limit. AMRIS, with its ORCA pipeline, T1–T7 decomposition, and iORCA cross-session calibration, is engineered specifically to externalise the chunks the human assessor cannot hold.
Three things change when observation is mechanised:
- Effective sample size per session approaches the demonstration opportunity count. Where the human assessor registers ~3.5 demonstrations per OB per session, the system registers all of them — plausibly 10–15 per OB depending on phase. The session becomes 3–4× more statistically informative.
- Cross-session pooling becomes automatic. Demonstrations persist as structured data and accumulate over the pilot’s entire training history. A four-session sequence, which would deliver ~14 demonstrations per OB to a human assessor without aggregation, delivers ~50 to a system with persistent state — well into Standard-tier per-OB reliability.
- Sampling bias is eliminated. The system does not preferentially attend to salient or expectation-confirming events. The sample is the population.
6.2 What does not change
The role of the human assessor does not disappear; it shifts. AMRIS does not grade pilots. It records, classifies, and surfaces evidence. The assessor calibrates the system, integrates contextual factors that no machine can read (a trainee’s evident anxiety; a brief lapse in attention with an obvious cause; the texture of a debrief), and exercises the judgment that turns evidence into a developmental conversation. The assessor is freed from the cognitively impossible task of real-time multi-stream observation, and reallocated to the cognitively appropriate task of integration, judgment, and trainee interaction.
This is a redistribution of labour aligned with the comparative advantages of human and machine. The machine does not approximate the assessor; the machine and the assessor each do the part of the assessment task they are equipped to do.
6.3 The regulatory argument
The framework specifies an assessment architecture; the framework does not specify an implementation. ICAO Doc 9995 is silent on whether the OBs must be observed by a human or by an instrumented system. EASA AMC1 ORO.FC.231 prescribes the competencies and the OBs but not the observational substrate. It is therefore open to operators and regulators to ask: given that the unaided human observational substrate cannot deliver the framework’s specified precision, what observational substrate can?
The argument we advance is not that AI-assisted observation is more efficient than human observation, although it is. It is that AI-assisted observation is the only known means of delivering the assessment fidelity the framework specifies. The arithmetic is in Sections 3 and 4. Without continuous observation and cross-session pooling, per-OB CBTA grading is statistically indefensible, and per-competency CBTA grading is reliable only at coarse precision. With them, both are reachable.
The question is not whether to deploy AI-assisted observation in CBTA. The question is how a CBTA programme without it can claim to deliver the assessment precision the framework specifies. The burden of proof has shifted.
6.4 What this means for AMRIS positioning
AMRIS is not a productivity tool that helps instructors do faster what they already do. AMRIS is a working-memory prosthetic and a statistical aggregation engine that, together, close the gap between the framework’s specifications and the human observer’s capacity. The right mental model is not “instructor-plus-software”; it is “instructor-plus-instrumentation” — the same model that has long been accepted in flight-data monitoring, where the FDR records and the human investigates.
The implication for the procurement conversation is that AMRIS is not competing against “instructor-only” assessment as an enhancement. It is competing against an assessment status quo whose internal validity, examined arithmetically, does not stand up. The defensible question for an airline is no longer “why would we buy this?” but “how do we currently defend the grades we issue?”
7. Conclusion
This paper has applied two well-established quantitative tools — Cowan’s working-memory model and Wilson’s binomial confidence interval — to the ICAO/EASA/IATA CBTA framework as documented. The findings are arithmetic, not opinion.
Instructor cognitive load during the highest-criticality phases of flight reaches 4.8× the established working-memory capacity limit. A single four-hour simulator session delivers ~3.5 demonstrations per OB, against a defensible-grade requirement of 25–97. Per-OB grading from a single session is, in formal statistical terms, indefensible. Per-competency grading is reliable at coarse precision tiers and only after multiple sessions are pooled.
These limits are not deficiencies of instructor skill. They are properties of the human cognitive architecture and of the binomial sampling distribution. They will not be addressed by better instructor training, by more granular grading rubrics, or by more careful framework documentation. They will be addressed only by changing the observational substrate.
AI-assisted observation, as instantiated in AMRIS, addresses both limits simultaneously. It externalises the chunks the human assessor cannot hold, and it accumulates the statistical sample sizes that human observation cannot. The result is not a new kind of assessment; it is the assessment the CBTA framework already specifies, finally delivered with the precision the framework requires.
Caveats and Future Work
Several modelling assumptions warrant explicit acknowledgement and invite further empirical refinement:
- Independence of demonstrations. The Wilson formula assumes Bernoulli independence between OB demonstrations. In practice, demonstrations are correlated within phase, within competency, and within crew. True required sample sizes are likely 1.5–2× the figures in Table 3.
- OB demonstration rate. The 1.5 demonstrations per active minute used in Section 4 is an order-of-magnitude estimate. Calibration against operational ORCA pipeline data — by phase, by aircraft type, by training programme phase — would tighten the per-tier session counts considerably. We invite operators with ORCA-instrumented training to share aggregated demonstration-rate data.
- Chunk-load model granularity. The chunk-load estimates in Section 3 use a structured but stylised decomposition (active sub-task, OBs in scope, TEM state, crew streams, competency tally). A more granular cognitive task analysis, ideally validated against eye-tracking or process-tracing studies of working CBTA assessors, would strengthen the cognitive-load argument.
- Word-picture compilation. We assume the five-point word-picture grade anchors are compiled into long-term memory for experienced assessors and excluded from the working-memory budget. This is conservative; for less experienced assessors the working-memory load is correspondingly higher.
- Inter-rater reliability. This paper addresses one source of measurement error — sample size. A complete reliability analysis would also examine inter-rater agreement (Cohen’s kappa, Krippendorff’s alpha) on identical observed material. Anecdotal evidence suggests inter-rater disagreement on CBTA OBs is substantial; quantifying it across a representative instructor population is a natural next study.
References
Adam, K. C. S., Mance, I., Fukuda, K., & Vogel, E. K. (2015). The contribution of attentional lapses to individual differences in visual working memory capacity. Journal of Cognitive Neuroscience, 27(8), 1601–1616.
Brown, L. D., Cai, T. T., & DasGupta, A. (2001). Interval estimation for a binomial proportion. Statistical Science, 16(2), 101–133.
Chase, W. G., & Simon, H. A. (1973). Perception in chess. Cognitive Psychology, 4(1), 55–81.
Cowan, N. (2001). The magical number 4 in short-term memory: A reconsideration of mental storage capacity. Behavioral and Brain Sciences, 24, 87–185.
Cowan, N. (2005). Working memory capacity. Hove: Psychology Press.
Cowan, N. (2010). The magical mystery four: How is working memory capacity limited, and why? Current Directions in Psychological Science, 19(1), 51–57.
EASA. (2022). AMC1 ORO.FC.231 Evidence-Based Training. European Union Aviation Safety Agency.
Ericsson, K. A., & Kintsch, W. (1995). Long-term working memory. Psychological Review, 102(2), 211–245.
Gilchrist, A. L., Cowan, N., & Naveh-Benjamin, M. (2008). Working memory capacity for spoken sentences decreases with adult ageing. Memory, 16(7), 773–787.
IATA. (2013). Evidence-Based Training Implementation Guide. International Air Transport Association.
IATA. (2024). Competency-Based Training and Assessment (CBTA) Expansion within the Aviation System. International Air Transport Association.
ICAO. (2013). Doc 9995: Manual of Evidence-Based Training. International Civil Aviation Organization.
Miller, G. A. (1956). The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological Review, 63(2), 81–97.
Rouder, J. N., Morey, R. D., Cowan, N., Zwilling, C. E., Morey, C. C., & Pratte, M. S. (2008). An assessment of fixed-capacity models of visual working memory. PNAS, 105(16), 5975–5979.
Shipstead, Z., Lindsey, D. R. B., Marshall, R. L., & Engle, R. W. (2014). The mechanisms of working memory capacity: Primary memory, secondary memory, and attention control. Journal of Memory and Language, 72, 116–141.
Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257–285.
Sweller, J., van Merriënboer, J. J. G., & Paas, F. (2019). Cognitive architecture and instructional design: 20 years later. Educational Psychology Review, 31, 261–292.
Wilson, E. B. (1927). Probable inference, the law of succession, and statistical inference. Journal of the American Statistical Association, 22(158), 209–212.