RL agents reveal what’s hard: Bootstrapping difficulty-ordered curricula for human learners
Publication Type
Conference Proceeding Article
Publication Date
7-2026
Abstract
Adaptive teaching algorithms require extensive human learning data, yet collecting such data is costly and time-intensive. We investigate whether reinforcement learning (RL) agents can serve as synthetic learners to bootstrap this process, warm-starting teacher algorithms before any human data is collected. We apply this framework to two teacher algorithms: PERM-H, which adjusts difficulty based on inferred learner ability, and SimMAC, which sequences tasks by jointly optimizing difficulty progression and inter-task similarity. Human studies with 464 participants across a platformer game and medical simulation show that RL-bootstrapped curricula significantly outperform random and control conditions while matching handcrafted sequences. Crucially, RL-derived difficulty estimates align with designer-intended difficulty, enabling effective easy-to-hard progressions without human data—and the failure of random training to improve over no training confirms that difficulty ordering, not mere exposure, drives learning gains.
Keywords
Adaptive Learning Systems, Cold-Start Problem, Curriculum Generation, Game-based Learning Environments, RL for Education
Discipline
Artificial Intelligence and Robotics
Research Areas
Intelligent Systems and Optimization
Publication
Proceedings of the 27th International Conference, AIED 2026, Seoul, South Korea, June 27 - July 3
First Page
424
Last Page
432
ISBN
9783032297549
Identifier
10.1007/978-3-032-29755-6_29
Publisher
Springer
City or Country
Cham
Citation
TIO, Sidney Xi Rong; LI, Wenjun; KARUNASENA, Galawala Ramesha Samurdhi; and VARAKANTHAM, Pradeep.
RL agents reveal what’s hard: Bootstrapping difficulty-ordered curricula for human learners. (2026). Proceedings of the 27th International Conference, AIED 2026, Seoul, South Korea, June 27 - July 3. 424-432.
Available at: https://ink.library.smu.edu.sg/sis_research/11259
Additional URL
https://doi.org/10.1007/978-3-032-29755-6_29