Brief №016 · August 2026

The Learning Agent Can Remove the Learning

AI can make a learner perform better while quietly performing part of the cognitive work that performance is supposed to demonstrate. The control problem is not whether to use AI assistance, but how to calibrate it, when to withdraw it, and what evidence counts as independent capability.

§ 01Bottom Line

AI assistance creates an easy measurement trap in learning: the system can make the learner look more capable before the learner is more capable.

A 2026 workplace-oriented experiment found that reliance on machine-learning predictions hindered the development of critical decision-making skills; when the aid became unavailable, participants who had relied on it experienced significant performance drops. 1 But the evidence does not support a simple “less AI is better” conclusion. Other 2026 studies found that AI-supported tutoring improved several forms of transfer and delayed performance when the design preserved verification, retrieval, and abstraction, and that hints without full solutions produced better transfer than hints plus solutions in a programming study. 23

The stronger conclusion is about calibration. If independent capability is the objective, assistance should be treated as part of the learning design: vary its depth, reduce it as proficiency grows, remove it before capability is inferred, and test whether performance persists after time and across changed contexts.

The best learning agent may not be the one that helps the most. It may be the one that knows when to stop helping.

§ 02Key Judgments
  1. Performance can rise faster than skill. In one experimental judgment task, reliance on machine-learning predictions hindered critical decision-making skill development and was followed by significant performance drops when the system was removed. 1
  2. More help is not automatically better for transfer. In a quasi-experimental programming study, hints plus full solutions produced stronger immediate-test gains, while hints alone produced better transfer-test performance. 3
  3. AI assistance can strengthen learning when it preserves the work of learning. A separate study found gains in near transfer, topic-shift transfer, and seven-day delayed performance, while broader transfer remained conditional on scaffolds that promoted verification, retrieval, and abstraction. 2
  4. Withdrawal is a measurement event, not merely the absence of a feature. If the organization needs independent capability, performance should be checked after assistance is removed and, where appropriate, after time has elapsed or the task has changed. 135
  5. Assisted and independent performance should not collapse into one signal. Where learning evidence informs consequential readiness or skill judgments, the assistance state should remain visible so the organization can distinguish human capability from human-plus-agent performance. 123
§ 03Analysis

The evidence supports a three-part interpretation. First, performance produced while an AI system is actively guiding, retrieving, correcting, or solving is partly a joint human-system output; it cannot automatically answer what the learner can do alone. Second, the shape of assistance matters: the studies in this evidence set show different immediate and transfer outcomes under different forms of support, while also showing that AI tutoring can improve learning when it preserves verification, retrieval, abstraction, and problem solving. Third, withdrawal is therefore part of measurement when independent capability is the objective. None of this means unaided performance is always the right goal or that more assistance is inherently harmful. It means the objective and the assistance regime must be explicit enough that the organization knows what its performance evidence actually represents.

Performance becomes a joint output.

Learning differs from ordinary productivity software because task completion is often only a proxy for the real objective. A learner may be expected to build judgment, retrieval fluency, error recognition, abstraction, or the ability to perform after support disappears. A system can improve the visible task while changing how much of that underlying work the learner actually performs.

The clearest evidence in this brief comes from The Dependency Dilemma: How Machine Learning Decision Aids can Undermine Skill Growth. In an experimental prediction-making task, the study found that reliance on machine-learning predictions hindered the development of critical decision-making skills and produced significant performance drops when the aid was unavailable. 1

That finding is study-bounded. It does not establish that every AI tutor creates dependency, or that enterprise learning agents are broadly weakening worker capability. What it does establish is a credible separation between performance with an aid and capability after the aid is removed.

That separation matters because high assisted performance is a joint human-system output. If the organization later uses that performance as evidence about the person alone, it needs another measurement step.

Help should preserve the work that learning requires.

The strongest counterevidence sharpens rather than weakens the argument.

A Frontiers study of AI-supported tutoring reported gains in near transfer, topic-shift transfer, and performance measured seven days later. The authors also concluded that broader transfer remained conditional on scaffolds that promoted verification, retrieval, and abstraction. 2

A quasi-experimental programming study found a different immediate-versus-transfer tradeoff. Learners receiving hints plus full solutions showed stronger improvement on immediate tests; learners receiving hints without the full solution performed better on the transfer test. The study linked the transfer result to deeper problem solving under a moderate cognitive load. 3

Together, these studies argue against treating assistance as a single variable. A full solution can be useful. So can a hint. The right design depends on the objective, the learner, the task, and the stakes. If the objective is immediate task completion, abundant assistance may be rational. If the objective is independent capability, some of the retrieval, judgment, and problem solving has to remain with the learner.

The design question is therefore not “How much can the agent do?” It is “Which work must the learner still do for the capability we care about to form?”

The control loop is incomplete until support is removed.

Autonoma’s synthesis is a five-part control model: calibrate, taper, delay, transfer, separate.

Calibrate assistance depth to the learner, task, and stakes instead of maximizing immediate success. Taper prompts, solutions, and agent initiative as proficiency develops. Delay at least some assessment until time has passed and the agent is absent. Transfer the capability into a changed context or authentic work sample. Separate assisted performance from evidence of independent capability.

The model is not a claim that one tapering schedule will work everywhere. It is a way to make the assistance itself observable and testable.

Professional-skills research makes that increasingly relevant. SocialCoach, a recent preprint, describes an LLM-powered agentic tutoring system for skills including negotiation and leadership and uses adaptive practice scheduling to personalize the learning journey. 4 Separately, research on AI tutoring for surgical skill acquisition assessed retention using a realistic surgical task and blinded technical-skill evaluation, illustrating the value of testing performance on authentic professional work rather than relying only on practice results. 5

Neither source proves enterprise-wide effectiveness or adoption. They do show why assistance calibration will matter more as AI tutoring moves closer to consequential professional skills.

Autonoma forecast: Within 18–24 months, leading enterprise learning platforms will begin exposing assistance-calibration controls such as configurable hint depth, progressive support withdrawal, delayed unaided checks, and separate reporting of assisted versus independent performance.

The timing is uncertain; the underlying product need is easier to see: once the agent materially affects the performance being measured, enterprises will need a way to distinguish the learner’s capability from the partnership’s output.

§ 04Indicators
  1. Assistance levels become explicit product primitives. Platforms distinguish hints, worked examples, and complete solutions instead of treating all AI help as one mode.
  2. Delayed checks appear after agent-supported practice. Programs assess learners after the agent is unavailable rather than measuring only supported performance.
  3. Transfer tasks sit beside practice scores. Assessment requires application in a changed context, not simply repetition of the trained task.
  4. Learning telemetry preserves assistance state. Performance is labeled assisted, partially assisted, or unaided.
  5. Adaptive tutors deliberately withdraw support. Systems require retrieval, judgment, or an attempted solution before escalating assistance.
  6. Professional simulations become capability evidence. AI-assisted practice is compared with authentic or unaided work samples before readiness is inferred.
§ 05Implications

For Chief Learning Officers.

Require separate measures of assisted performance, delayed retention, transfer, and unaided capability when a program is meant to establish independent readiness. A high score achieved with continuous agent support should not carry more meaning than the measurement design can defend.

For instructional-design and assessment leaders.

Build an assistance ladder, not merely an AI feature. Decide what the learner should retrieve, judge, generate, or correct before the agent intervenes, and reduce assistance as those capabilities develop.

For learning-platform and AI-tutor product owners.

Make assistance depth observable. Expose when the system supplied a hint, worked example, partial solution, or complete solution; make tapering configurable; preserve that provenance in learning telemetry.

For certification and compliance owners.

If a credential or readiness decision implies independent performance, require an independent performance check. Agent-assisted practice can be useful evidence without being sufficient evidence.

For skills and workforce leaders.

Do not convert AI-assisted learning telemetry directly into a durable skill record unless the record preserves assistance state or is backed by an independent capability check.

For employees and learners.

Make the rules legible. People should know when assistance is being measured, when they will be expected to perform independently, and when the system will deliberately hold back an answer so they can attempt the work themselves.

§ 06Dissenting View

Weight: Strong. The strongest counterargument is that the brief risks making assistance itself look suspicious when good tutoring has always involved scaffolding, examples, feedback, and temporary support.

The evidence supports that objection. AI-supported tutoring in the Frontiers study improved several forms of transfer and delayed performance. In the programming study, full solutions were associated with stronger immediate-test improvement even though hints alone produced stronger transfer. 23 Different learners, skills, and stakes will require different assistance levels.

The central dependency evidence also comes from a workplace decision-aid experiment, not an enterprise learning-agent deployment. 1 That is a meaningful limit on generalization.

The appropriate conclusion is therefore not “reduce AI assistance.” It is narrower: do not infer independent capability from assisted success without testing whether the capability survives the assistance. If an organization can show durable retention, transfer, and authentic performance after support is removed, the central risk described here is materially reduced for that use case.

§ NoteThe Architect’s Note

Enterprise learning has spent years trying to move beyond completion as the measure of capability. Agentic AI introduces a harder version of the same problem: a learner can now produce an impressive performance while an intelligent system is explaining, retrieving, correcting, and deciding alongside them.

That may be valuable performance support. It may also be excellent instruction. But if the capability disappears with the agent, the enterprise has measured the partnership and recorded it as the person.

The design objective should not be to remove the agent. It should be to know what remains when it leaves.

Methodology

This brief draws on five public research sources spanning a workplace decision-aid experiment, learning-transfer studies, an emerging professional-skills tutoring architecture, and professional-task assessment. Each empirical claim is kept within its study setting; the SocialCoach preprint is used only as architecture and use-case evidence, not as proof of effectiveness or adoption. The five-part control model and the 18–24 month platform forecast are Autonoma Intelligence synthesis rather than findings attributed to a single study.

Sources

  1. The Dependency Dilemma: How Machine Learning Decision Aids can Undermine Skill Growth. Peer-reviewed workplace decision-aid research, Springer, 2026.
  2. The causal effects of artificial intelligence use on metacognition, engagement, and knowledge transfer in educational contexts. Frontiers in Psychology, 2026.
  3. Balancing error-correction hints and solution guidance: a quasi-experimental study of ChatGPT-integrated feedback strategies in Jupyter-based programming education. Nature Portfolio, 2026.
  4. SocialCoach: An LLM-Powered Agentic Tutoring System for Personalized Social Skill Development. arXiv preprint, 2026.
  5. AI tutoring versus expert human instruction for surgical skill acquisition. Peer-reviewed professional-training research, Springer/BMC, 2026.