RL agents reveal what’s hard: Bootstrapping difficulty-ordered curricula for human learners

Publication Type

Conference Proceeding Article

Publication Date

7-2026

Abstract

Adaptive teaching algorithms require extensive human learning data, yet collecting such data is costly and time-intensive. We investigate whether reinforcement learning (RL) agents can serve as synthetic learners to bootstrap this process, warm-starting teacher algorithms before any human data is collected. We apply this framework to two teacher algorithms: PERM-H, which adjusts difficulty based on inferred learner ability, and SimMAC, which sequences tasks by jointly optimizing difficulty progression and inter-task similarity. Human studies with 464 participants across a platformer game and medical simulation show that RL-bootstrapped curricula significantly outperform random and control conditions while matching handcrafted sequences. Crucially, RL-derived difficulty estimates align with designer-intended difficulty, enabling effective easy-to-hard progressions without human data—and the failure of random training to improve over no training confirms that difficulty ordering, not mere exposure, drives learning gains.

Keywords

Adaptive Learning Systems, Cold-Start Problem, Curriculum Generation, Game-based Learning Environments, RL for Education

Discipline

Artificial Intelligence and Robotics

Research Areas

Intelligent Systems and Optimization

Publication

Proceedings of the 27th International Conference, AIED 2026, Seoul, South Korea, June 27 - July 3

First Page

424

Last Page

432

ISBN

9783032297549

Identifier

10.1007/978-3-032-29755-6_29

Publisher

Springer

City or Country

Cham

Additional URL

https://doi.org/10.1007/978-3-032-29755-6_29

This document is currently not available here.

Share

COinS