Audit Packet: The Learning Agent Can Remove the Learning
What the evidence supports, how the assistance-calibration thesis was tested, and where the argument remains deliberately bounded.
This audit packet supports Brief №016: The Learning Agent Can Remove the Learning. Read the brief first for the full argument.
Autonoma briefs are designed to be inspectable. This packet shows what the brief claims, how each claim was tested, what it does not claim, and where caveats remain — without exposing raw internal logs, prompts, operator notes, source-routing mechanics, hashes, local paths, secrets, or unpublished candidate claims.
← Open Brief №016 — The Learning Agent Can Remove the Learning
Brief Summary and Audit Verdict
The central thesis, what the evidence supports, and what it does not.
Evidence verdict: SUPPORTED WITH MATERIAL CAVEATS.
The central thesis is supported: assisted performance and independent post-assistance capability can diverge, and the design of the assistance can influence what transfers after practice.
The strongest direct evidence comes from a workplace-oriented experimental decision task in which reliance on machine-learning predictions hindered critical decision-making skill development and was followed by significant performance drops when the aid became unavailable. 1 The evidence also contains strong positive counterevidence: AI-supported tutoring produced gains in several forms of transfer and delayed performance when scaffolding preserved verification, retrieval, and abstraction, while a programming study found stronger transfer under hints-only feedback than under hints plus full solutions. 23
The evidence therefore supports an assistance-calibration thesis, not an anti-assistance thesis.
It does not establish that all AI assistance harms learning, that enterprise learning agents are currently reducing workforce capability at scale, that hints are universally superior to solutions, that one tapering schedule works across domains, or that the professional-skills systems described here have proven enterprise effectiveness. The 18–24 month platform forecast remains a forecast, not an observed market fact.
Claim Register
Each public claim, its evidence, the verdict, and the boundary.
| Public claim | Evidence | Verdict | Boundary |
|---|---|---|---|
| Reliance on machine-learning predictions hindered critical decision-making skill development in an experimental judgment task. | 1 | Supported | One workplace-oriented experimental task; not enterprise learning-agent prevalence. |
| Participants who relied on the aid experienced significant performance drops when it became unavailable. | 1 | Supported | Study-bounded withdrawal effect; not proof that every form of AI support creates dependency. |
| AI-supported tutoring produced gains in near transfer, topic-shift transfer, and seven-day delayed performance. | 2 | Supported | Education context; broader transfer remained conditional. |
| Broader transfer remained conditional on scaffolds promoting verification, retrieval, and abstraction. | 2 | Supported within study scope | Does not establish one universal scaffold design. |
| Hints-only feedback produced better transfer-test performance than hints plus full solutions in a programming study. | 3 | Supported | Programming education; not a universal assistance hierarchy. |
| Hints plus full solutions produced stronger immediate-test improvement than hints-only feedback. | 3 | Supported | Shows an immediate-versus-transfer tradeoff within this study. |
| Agentic AI tutoring is being explored for professional skills such as negotiation and leadership using adaptive practice scheduling. | 4 | Supported as architecture/use-case evidence | Preprint; not outcome, adoption, or enterprise-effectiveness proof. |
| Realistic professional tasks can be used to assess learning and skill retention rather than relying only on practice performance. | 5 | Supported as assessment-context evidence | Does not directly test assistance withdrawal. |
| Assisted and independent performance should be measured separately when independent skill matters. | 1235 | Autonoma analytic synthesis | Governance recommendation, not a direct finding from one source. |
| Enterprise learning platforms should preserve assistance state and support calibrated withdrawal. | 1234 | Autonoma analytic synthesis | Design recommendation. |
| Leading enterprise learning platforms will add assistance-calibration controls within 18–24 months. | 1234 | Autonoma forecast | Product adoption and timing are not established facts. |
The brief intentionally does not translate study coefficients into generalized enterprise effect sizes. The source set does not support that extrapolation.
Source Ledger
What each source is competent to prove — and its material limitation.
| Source | Source class | Role in brief | Material limitation |
|---|---|---|---|
| The Dependency Dilemma 1 | Peer-reviewed workplace decision-aid study | Central dependency and withdrawal mechanism | One experimental judgment task; not an enterprise learning-agent deployment |
| The causal effects of AI use… 2 | Peer-reviewed learning-transfer study | Positive counterevidence; transfer and scaffolding | Education context limits enterprise generalization |
| Balancing error-correction hints… 3 | Peer-reviewed quasi-experimental feedback study | Assistance-calibration evidence | Programming education; not a universal tapering schedule |
| SocialCoach 4 | Recent independent preprint | Professional-skills bridge; agentic tutor architecture | Does not establish effectiveness or enterprise adoption |
| AI tutoring vs. expert human instruction (surgical) 5 | Peer-reviewed professional-training research | Authentic professional-task assessment context | Does not directly establish an assistance-withdrawal effect |
The central mechanism rests on independent empirical research rather than vendor claims. SocialCoach is used only to establish an emerging professional-skills architecture with adaptive scheduling; it is not used as evidence of learning outcomes. The surgical-training study is similarly bounded to authentic professional-task assessment context.
Evidence Boundaries
The four bounded conclusions, and the precise editorial boundary.
The evidence supports four bounded conclusions. First, AI assistance can improve immediate performance while independent skill formation follows a different trajectory. 13 Second, removal of an aid can expose a gap between supported and independent performance. 1 Third, AI tutoring can improve transfer and delayed performance when the design preserves cognitive work such as verification, retrieval, and abstraction. 2 Fourth, different assistance levels can create different immediate-versus-transfer tradeoffs. 3
The evidence does not justify an estimate of how frequently enterprise learning agents create dependency; a claim that current platforms misclassify assisted performance at scale; a universal causal claim covering all AI tutors, workers, or tasks; a prescribed numerical tapering schedule; a claim that cognitive struggle is always beneficial; a named-vendor accusation; a quantified productivity or workforce impact; or a regulatory requirement for assistance withdrawal.
The central editorial boundary is therefore precise:
Successful performance while an agent is helping does not, by itself, establish capability after that help is withdrawn.
That is materially different from claiming that AI assistance prevents learning.
Adversarial Review and Falsification
The strongest objections, their weight, and how each could be falsified.
Counterargument 1: Well-designed AI tutoring can improve learning and transfer.
Weight: Strong. Accepted.
The Frontiers study reported gains in near transfer, topic-shift transfer, and delayed performance under AI-supported tutoring. 2 The brief therefore cannot recommend reducing assistance by default.
Falsification test: Compare delayed, unaided, and new-context performance between learners receiving calibrated AI assistance and a credible alternative. If assisted learners retain and transfer the capability after the AI is removed, the dependency concern is materially weakened for that implementation.
Counterargument 2: Full solutions may be better for some learners and objectives.
Weight: Strong. Accepted.
In the programming study, hints plus full solutions produced stronger immediate-test improvement even though hints alone produced stronger transfer. 3 A novice, a time-critical task, an accessibility need, or a performance-support use case may rationally favor more assistance.
Falsification test: Predefine whether the objective is immediate completion, learning, transfer, or independent certification, then test assistance levels against that objective rather than assuming one mode is always superior.
Counterargument 3: The central dependency evidence is not a learning-agent deployment.
Weight: Strong. Accepted.
The strongest direct evidence comes from a workplace decision-aid experiment. 1 Applying the mechanism to enterprise learning agents is an analytic bridge.
Falsification test: In an enterprise-like learning task, compare sustained high assistance with calibrated or tapered assistance and measure immediate performance, delayed performance, transfer, and unaided execution. Failure to reproduce a meaningful post-withdrawal difference would lower confidence in the enterprise-learning extension.
Counterargument 4: Learning effects are too context-dependent for a platform-level control model.
Weight: Moderate. Partly accepted.
Learner expertise, task complexity, domain, feedback design, and assessment timing can all change the value of assistance. The evidence does not establish one optimal assistance schedule. That does not eliminate the need for control. It means the control should expose assistance depth and make it adjustable rather than encoding one universal rule.
Falsification test: If assistance depth shows no meaningful relationship to delayed or transfer performance across multiple credible domains, the case for assistance-calibration controls would weaken.
Editorial Decisions
Judgment calls made in shaping the brief.
- Calibrate assistance, do not stigmatize it. Positive AI-tutoring evidence is part of the core argument rather than being buried as a caveat. 23
- Separate performance from capability throughout. The brief avoids treating a successful assisted task as a learning outcome unless the evidence supports persistence or transfer.
- Keep the workplace evidence bounded. The decision-aid study supplies a mechanism and causal result within its experimental setting; it does not justify a prevalence claim. 1
- Use the professional-skills bridge narrowly. SocialCoach supports the existence of an agentic tutoring architecture for professional social skills with adaptive scheduling, not an effectiveness claim. 4
- Use the surgical study as assessment context only. It supports realistic professional-task assessment as a retention measure, not an assistance-withdrawal effect. 5
- Label the control model as synthesis. Calibrate → Taper → Delay → Transfer → Separate is Autonoma’s design synthesis across the evidence set.
- Keep the forecast separate from observed facts. The 18–24 month platform prediction does not imply that those controls are common today.
Reader-Facing Caveats
The limits that narrow the claim.
- The central workplace evidence comes from one experimental judgment task.
- Important positive-control evidence comes from educational settings rather than enterprise workforce deployments.
- The programming study does not establish that hints-only assistance is superior for every task or learner.
- SocialCoach is a recent preprint and is used only as architecture and use-case evidence.
- The evidence set does not estimate current enterprise adoption or prevalence.
- Assisted-versus-unaided capability is likely to vary by skill type, learner, context, and stakes.
- The 18–24 month product forecast is an editorial forecast, not an observed market fact.
These limitations narrow the claim; they do not erase the measurement problem the brief asks enterprises to test.
Correction Log
Post-publication corrections, if any.
No corrections as of initial publication.
Editorial Signoff
The human editorial judgment on readiness.
The evidence base is sufficient to support the central assistance-calibration thesis with the boundaries documented in this packet. All five load-bearing public evidence claims used in the brief are supported by the source record, and the strongest counterevidence materially shapes the recommendation from “less assistance” to “better-calibrated assistance.”
Final human editorial approval: Approved.
A Reproducible Enterprise Test
A sandboxed assisted-vs-unaided capability test using synthetic tasks.
An organization can test the central Brief 016 mechanism in a sandbox without using real employee decisions.
Choose one bounded skill with an observable work product—for example, interpreting a fictional business scenario, troubleshooting a simulated system fault, writing a structured recommendation, completing a synthetic negotiation, or performing a sandboxed technical procedure. Use synthetic or training-only tasks.
Compare at least three assistance conditions
- High assistance: proactive guidance, worked examples, and complete solutions when requested.
- Calibrated assistance: hints first; retrieval or an attempted solution before escalation; progressively less support across successful repetitions.
- Minimal assistance: reference material and conventional feedback without generative solution support during practice.
Measure four outcomes separately
- Assisted performance: performance while the assigned support is present.
- Delayed retention: performance after time has elapsed.
- Transfer: performance on a materially changed problem.
- Unaided performance: performance when the AI is absent.
For every measured performance, record whether it was fully assisted, partially assisted, or unaided, and record the type of help used—hint, worked example, retrieval support, error correction, partial solution, or complete solution.
The stronger pass condition is not simply a high practice score. It is that learners maintain acceptable performance after assistance is removed, after time has elapsed, and when the task changes enough to require transfer.
If high-assisted practice produces the best immediate results and equivalent or better delayed, transfer, and unaided performance, abundant assistance may be appropriate for that skill. If immediate performance rises while delayed or unaided performance falls, the assistance has become part of the performance being measured and should be represented accordingly.
Methodology
How sources were reviewed and what was deliberately excluded.
This audit evaluated five public research sources against the brief’s public claims and kept each claim within the scope of the underlying evidence. Empirical findings, Autonoma analytic synthesis, and the 18–24 month forecast are treated as separate evidence classes. No claim about enterprise prevalence, named-vendor outcomes, regulatory mandate, or quantified business impact is required for the brief’s thesis.
Sources
Numbered to match the citations in Brief №016 and in this packet.
- The Dependency Dilemma: How Machine Learning Decision Aids can Undermine Skill Growth. Peer-reviewed workplace decision-aid research, Springer, 2026.
- The causal effects of artificial intelligence use on metacognition, engagement, and knowledge transfer in educational contexts. Frontiers in Psychology, 2026.
- Balancing error-correction hints and solution guidance: a quasi-experimental study of ChatGPT-integrated feedback strategies in Jupyter-based programming education. Nature Portfolio, 2026.
- SocialCoach: An LLM-Powered Agentic Tutoring System for Personalized Social Skill Development. arXiv preprint, 2026.
- AI tutoring versus expert human instruction for surgical skill acquisition. Peer-reviewed professional-training research, Springer/BMC, 2026.