Research note

Who Teaches
the Learner?

A framework-first review of curriculum as an outer decision problem: what the teacher observes, controls, optimizes, and chooses not to model.

Daoyuan GuoAugust 202617 min read

A curriculum is not only a sequence of data. It is a sequence of interventions on a changing learner.

Suppose a learner can study algebra or geometry. Geometry produces the larger improvement today, so \(\Delta J_t(\text{geometry})>\Delta J_t(\text{algebra})\). But algebra may be a prerequisite that makes geometry dramatically more valuable tomorrow. Then

\[\text{best immediate lesson}\neq\text{best curriculum decision}.\]

This small example contains most of the difficulty. Curriculum is not synonymous with easy-to-hard ordering. A lesson intervenes on a learner whose capabilities change because of that intervention; once the selection rule adapts from feedback, curriculum design becomes an outer decision problem.

This is not a survey of curriculum learning as a whole. Fixed easy-to-hard schedules, self-paced learning, competence rules, and pretraining-data curricula are important neighboring ideas. I focus on curriculum mechanisms viewed as adaptive outer decision processes.

For this note, I use adaptive curriculum optimization as a descriptive umbrella. Within it, RL for Curriculum refers more narrowly to bandit, contextual-decision, and MDP/POMDP teachers that learn a curriculum decision rule from feedback. The organizing question is: what does the teacher observe, what can it change, how is usefulness measured, and how much of the resulting dynamics should it model?

\[\boxed{\text{true learner dynamics}\rightarrow\text{partial observations and proxies}\rightarrow\text{algorithmic approximation}\rightarrow\text{curriculum decision}}\]

1. The teacher's decision problem

Latent student state and teacher observation

Let \(z_t\) denote the student's latent learning state. This is conceptual, not necessarily the parameter vector \(\theta_t\). It may include skill mastery, forgotten knowledge, readiness for transfer, optimizer history, interference between tasks, and which skills remain learnable.

The teacher rarely observes \(z_t\) directly. It sees an observation constructed from training history:

\[o_t^C\sim\mathsf O^C(\cdot\mid z_t,h_t),\qquad\text{or simply}\qquad o_t^C=\Phi(h_t).\]

The history \(h_t\) may contain recent losses, evaluations, solve rates, advantages, probe performance, prior curriculum choices, gradient statistics, or capability summaries. A computed learner summary is an observation of student state; it is not automatically a Markov state.

Teacher action and student update

A curriculum action is an intervention on the student's training experience:

\[a_t^C=\text{task, source, difficulty range, category, cluster, batch, problem, goal, or environment}.\]

The teacher usually does not edit the student's parameters directly. It changes the data, task, or environment from which the next update is produced:

\[\theta_{t+1}=\mathcal A\!\left(\theta_t,\mathcal B_t(a_t^C)\right).\]

Teacher transition, reward, and objective

At the conceptual level, the intervention changes latent student state:

\[z_{t+1}\sim P^C(\cdot\mid z_t,a_t^C),\qquad a_t^C\rightarrow z_{t+1}\rightarrow\text{future curriculum utility}.\]

An ideal teacher signal may depend on the before-state, intervention, and after-state, \(r_t^C=G(z_t,a_t^C,z_{t+1})\). Its objective might be final student quality or cumulative outer return:

\[\max_{\pi^C}\mathbb E[J(\theta_T)]\qquad\text{or}\qquad\max_{\pi^C}\mathbb E\!\left[\sum_{t=0}^{T-1}\gamma_C^t r_t^C\right].\]
\[\boxed{J^C\longrightarrow R^{C,*}\longrightarrow\widehat R^C}\]

Here \(J^C\) is the final training objective; \(R^{C,*}\), when it is meaningful, is a local credit decomposition consistent with that objective; and \(\widehat R^C\) is the signal a method can actually measure. Not every curriculum objective admits a uniquely meaningful one-step teacher reward. Most systems therefore construct a proxy—learning progress, validation gain, solve rate, advantage magnitude, or estimated policy improvement—and that instrumentation choice is central.

latent student→teacher observation→intervention→student update→proxy signal

A terminology boundary

\[\boxed{\begin{aligned}\text{Curriculum for RL}&=\text{curriculum mechanism}+\text{RL student},\\\text{RL for Curriculum}&=\text{feedback-learned outer decision rule}+\text{student learner}.\end{aligned}}\]

FastCuRL, for example, uses staged data segmentation and context extension around an RL-trained reasoning model; its outer schedule is designed rather than learned online. DUMP, SEC, and SPaCe instead adapt the data-selection mechanism from training feedback. The inner optimizer being GRPO does not decide the category—the outer mechanism does.

Underlying process and method approximation

The cleanest separation is between a reference teacher–student process we would ideally like to reason about and the abstraction a concrete method actually implements:

\[\boxed{\mathcal E^C=(\mathcal Z^C,\mathsf O^C,\mathcal A^C,P^C,J^C)}\]
\[\begin{aligned}\mathcal Z^C&:\text{ latent learner state},&\mathsf O^C&:\text{ observation kernel},&\mathcal A^C&:\text{ interventions},\\P^C&:\text{ learner dynamics},&J^C&:\text{ ultimate training objective}.&&\end{aligned}\]

A concrete method implements an approximation of that process:

\[\boxed{\mathcal M^C=(\widehat{\mathcal O}^C,\widehat{\mathcal A}^C,\widehat R^C,\widehat P^C,\Pi^C,H^C,\kappa^C)}\]
\[\begin{aligned}\widehat{\mathcal O}^C&:\text{ student information used},&\widehat{\mathcal A}^C&:\text{ action abstraction exposed},\\\widehat R^C&:\text{ observable curriculum signal},&\widehat P^C&:\text{ assumed or implicit dynamics},\\\Pi^C&:\text{ decision-update rule},&H^C&:\text{ effective planning horizon},&\kappa^C&:\text{ information and decision cost}.\end{aligned}\]

In this note, I treat each curriculum method as an approximation to a common reference process \(\mathcal E^C\), implemented through a method-specific abstraction \(\mathcal M^C\). The separation between the process we would ideally reason about and the information and structure an algorithm actually models is the review's organizing device.

2. Why curriculum is a difficult outer-loop learning problem

The six difficulties below belong to three groups: dynamics, information, and statistics/systems.

\[\boxed{\begin{array}{lll}\textbf{Dynamics}&\text{sequentiality}&\text{endogenous non-stationarity}\\[2mm]\textbf{Information}&\text{partial observability}&\text{proxy reward}\\[2mm]\textbf{Statistics / systems}&\text{expensive feedback}&\text{large action spaces}\end{array}}\]

2.1 Sequentiality

Immediate usefulness \(\mathbb E[r_t^C\mid z_t,a_t^C]\) and long-term curriculum value \(Q^C(z_t,a_t^C)\) need not agree. Algebra can enable geometry; an easy example can build a representation used much later; revisiting a task can prevent forgetting; one domain can transfer to another.

\[a_t^C\rightarrow z_{t+1}\rightarrow z_{t+2}\rightarrow\cdots\]

The curriculum problem becomes genuinely sequential when today's lesson changes what can be learned efficiently tomorrow.

2.2 Endogenous non-stationarity

Sequentiality concerns the value of delayed consequences; endogenous non-stationarity concerns how today's intervention changes the local utility landscape the teacher will face later.

A non-stationary bandit may track \(\mu_k(t)=\mathbb E[r_t^C\mid a_t^C=k]\). In curriculum learning, however, drift is often endogenous:

\[a_t^C\rightarrow z_{t+1}\rightarrow\mu_k(t+1).\]

The teacher is not merely tracking a changing environment; its previous interventions help create the changes it later adapts to. Non-stationary bandits can track moving arm utility without explicitly modeling how today's action causally changes tomorrow's arms. That is a simplification, not a defect.

\[\boxed{\text{tracking }\mu_k(t)\neq\text{modeling how }a_t^C\text{ changes }\mu_k(t+1)}\]

Curriculum non-stationarity is often endogenous: the teacher changes the learner that generates future arm values. This is what a utility-tracking approximation leaves implicit.

This view is now explicit in recent LLM curriculum work. Manifold Bandits treats problem sampling as a structured bandit with endogenous non-stationarity: selected prompts affect how learning signals evolve over the policy's latent task geometry.

2.3 Partial observability

The teacher sees \(o_t^C=\Phi(h_t)\), not the full learning state. Two students with the same recent loss can differ in mastery, transfer potential, forgotten capabilities, or gradient geometry. A POMDP is therefore a natural reference abstraction whenever no available learner summary is sufficient for predicting curriculum consequences. The practical question is not “do we have a state vector?” but “is this student representation sufficiently informative for curriculum decisions?”

2.4 Proxy reward and myopia

Two problems should be kept separate. Proxy alignment asks whether \(\widehat r_t^C\) predicts desirable final performance \(J(\theta_T)\). Planning asks whether maximizing even a correct immediate gain yields the best sequence. The second problem remains when

\[\widehat r_t^C=J(\theta_{t+1})-J(\theta_t).\]

A target-aligned one-step estimate does not remove prerequisite, transfer, or delayed-effect structure.

2.5 Expensive outer feedback

select data→train student→roll out / evaluate→measure utility

One curriculum transition can contain many inner training samples. Especially for LLM post-training, \(N_{\text{outer decisions}}\ll N_{\text{inner samples}}\). A rich outer model may be expressive yet impossible to estimate from the available teacher transitions. Teacher sample efficiency is part of the problem definition.

2.6 Large and structured action spaces

The teacher may choose among twelve tasks or millions of problems. Large action spaces raise exploration cost and create sparse feedback, but also contain structure: categories, clusters, embeddings, hierarchies, and generative parameterizations can share evidence across related interventions. The resolution at which a teacher acts is a statistical modeling choice.

The literature can now be read as a collection of approximations from \(\mathcal E^C\) to \(\mathcal M^C\). Some methods compress student information; others simplify the intervention space, proxy, transition structure, horizon, or cost. This two-level framework is my organizational device, not standard notation.

3. How has RL for Curriculum been modeled?

The table gives the map; the subsections that follow extract the recurring modeling choices behind it.

MethodEra / domainTeacher observationCurriculum actionTeacher signalOuter optimizerOuter temporal model
Graves et al.2017, supervisedTask reward history; no explicit student stateTask / data sourcePrediction or complexity gainExp3.SBandit; no transition model
Matiisen et al.2017, supervised / RLTask performance slopesSubtaskLearning progress, including forgettingBandit-style teacherReactive progress tracking
Learning to Teach2018, supervisedStudent, data, and history featuresRetain / discard exampleTime to target accuracyREINFORCEDiscounted teaching episode
Narvekar & Stone2019, RL transferRepresentation of agent knowledgeSource task / transferNegative time-to-learnSARSA(\(\lambda\))Explicit curriculum MDP
El-Bouri et al.2020, supervisedEncoded student weightsBatch or continuous data regionValidation-weighted progressDQN / DDPGDiscounted teaching episode
PAIRED2020, environment designPartially built environmentConstruct environmentProtagonist–antagonist return gapPPOSequential environment construction
FastCuRL2025, LLM RFT — boundary caseFixed stage indexData segment / context stageDesigned scheduleFixed schedulerPredefined stage sequence
DUMP2025, LLM RFTRecent distribution advantages and countsDistribution mixtureMean absolute advantageUCBNon-stationary arm utilities
SEC2025, LLM RFTCategory valuesReasoning categoryMean absolute advantageBoltzmann + TD(0)Bootstrapped category utilities
SPaCe2025/26, LLM RFTCluster solve rate, pulls, stagnationData clusterCorrectness / hardness / progressThompson SamplingNon-stationary cluster utilities
Manifold Bandits / BMC2026, LLM RFTPrompt-level learning-signal beliefs + latent task treeHierarchically related promptsRollout-reward variabilityHierarchical Thompson SamplingStructured bandit with endogenous drift
Actor–Curator2026, LLM RFTCandidate problems; online utilitiesProblem probabilitiesPer-problem policy-improvement estimateNeural OSMD with clippingBandit over one-update improvement

3.1 What does the teacher know about the student?

Earlier work explores three broad strategies: no explicit state, hand-designed behavioral summaries, and learned embeddings. Graves represents student change only through task reward history; Kumar uses likelihoods on prototype sentences as a compact behavioral summary; Reinforcement Teaching learns an embedding from probe behavior. Other state-conditioned teachers use student, data, history, knowledge, or parameter features—the table records the distinctions.

Recent LLM schedulers often avoid a compact Markov representation of the whole student. Instead, learner evolution is absorbed into moving distribution-, category-, cluster-, or problem-level utility estimates. DUMP's advantage windows and SPaCe's solve-rate statistics illustrate this alternative.

\[\boxed{\text{student-state conditioning}\qquad\text{and}\qquad\text{utility tracking}}\]

These are recurring strategies, not mutually exclusive ends of a continuum. A method can condition on student features while also tracking non-stationary utilities. BMC illustrates a third modeling choice: it leaves student state implicit but represents structure among the actions themselves.

3.2 What can the teacher choose?

Across pre-LLM work, actions include tasks or sources (Graves; Matiisen), data/noise bins (Kumar), batches or continuous regions of difficulty-sorted data (El-Bouri), source tasks and transfer choices (Narvekar and Stone), and constructed environments (PAIRED). At LLM scale, DUMP acts on a distribution mixture, SEC selects reasoning categories, SPaCe selects semantic/difficulty clusters, and Actor–Curator assigns selection probabilities to individual candidate problems.

\[\text{coarse shared action}\quad\longleftrightarrow\quad\text{fine local action}\]
\[\text{independent actions}\quad\longleftrightarrow\quad\text{structured, related actions}\]

Granularity and structure are separate axes. Coarse actions share evidence by construction; fine actions offer precise intervention but demand far more feedback unless relations between actions can be exploited. Actor–Curator generalizes across problem representations, while BMC builds a hierarchical task tree from policy embeddings so evidence can propagate among related prompts. Thus \(\widehat{\mathcal A}^C\) specifies both intervention resolution and what cross-action structure the teacher can use.

3.3 What counts as a useful lesson?

Learning progress. Graves et al. derive rewards from prediction gain or changes in model complexity. Matiisen et al. use task-performance slopes, so decreasing performance can trigger review. Reinforcement Teaching also uses learning-progress shaping. These signals try to locate the learner's moving frontier:

\[\widehat r_t^C\approx m_{t+1}-m_t.\]

Training efficiency. Learning to Teach uses a terminal time-to-threshold signal in its data-teaching experiment; Narvekar and Stone assign negative time-to-learn as curriculum cost. Here the lesson is useful because it reduces training effort.

Regret-like utility. PAIRED trains an environment designer against a protagonist and antagonist, using their return gap as a regret approximation. It seeks environments that expose weaknesses while remaining solvable by another agent; this is a specific unsupervised environment-design construction, not a generic curriculum reward.

Inner-learning signals at LLM scale. DUMP and SEC use expected absolute advantage as a proxy for immediate learnability:

\[\widehat r_t^C\approx\mathbb E[|\widehat A|].\]

It indicates the presence of immediate policy-learning signal already available inside RL training. It does not directly measure long-term curriculum value. SPaCe instead uses online binary correctness aggregated into cluster solve rates, converts solve rate to hardness, and discounts persistently stagnant clusters.

Policy-improvement estimates. Actor–Curator estimates per-problem contributions to \(J(\pi_{t+1})-J(\pi_t)\) from pre- and post-update policies. This is a first-order credit-assignment signal for the performance gain induced by one actor update; it does not by itself solve long-horizon curriculum planning.

\[\boxed{\text{signal semantics}\quad\times\quad\text{measurement cost}}\]

Learning progress, training efficiency, regret, inner-optimization signal, and policy improvement describe different semantics—not a monotonic progression. Each signal has its own cost, variance, alignment, and horizon; expensive measurement is not automatically better aligned. BMC adds a useful diagnostic question: is the instrument measuring productivity (learning signal), diversity (task-manifold coverage), or utility (evaluation relevance)? These aspects of curriculum value are not interchangeable.

3.4 How does the teacher learn and adapt?

A fixed scheduler, such as FastCuRL's staged segmentation and context extension, specifies \(\pi_t^C\) in advance. Bandit teachers instead learn action utilities: Graves uses non-stationary Exp3.S; DUMP uses UCB with recent advantage windows; SEC combines Boltzmann sampling with TD(0) category-value updates; SPaCe uses progress-aware Thompson Sampling; Actor–Curator uses neural function approximation with an OSMD-derived, PPO-style clipped curator objective. BMC uses hierarchical Thompson Sampling over a latent task tree, combining Bayesian filtering with bottom-up empirical-Bayes propagation so evidence can be shared across related prompts.

Contextual or state-conditioned teachers choose \(a_t^C\sim\pi^C(\cdot\mid o_t^C)\). Learning to Teach uses REINFORCE on student/data/history features. Kumar uses DQN on prototype-sentence likelihoods. El-Bouri uses DDPG for continuous batch location and width, and DQN for a discrete variant. Narvekar and Stone learn a curriculum-MDP policy with SARSA(\(\lambda\)). PAIRED uses PPO for its environment-design agents. Reinforcement Teaching uses DQN/Double DQN teachers with learned behavioral state embeddings.

These algorithms attempt to exploit state dependence and, in some cases, delayed effects. Their use of an RL optimizer does not guarantee that the observation is Markov or that useful long-horizon planning is learned.

Explicit-state adaptation changes the action because \(o_t^C\) changes: \(a_t^C\sim\pi^C(\cdot\mid o_t^C)\). Implicit utility tracking changes the action because estimates \(\widehat\mu_k(t)\) move. Graves shares weight across a non-stationary bandit; DUMP uses a sliding window; SEC bootstraps category values with TD(0); SPaCe updates solve-rate and stagnation statistics; Actor–Curator receives online partial-feedback estimates after actor updates.

Explicit student-state modeling and non-stationary arm tracking are two different ways to adapt to a changing learner. Neither automatically models the causal effect of today's action on tomorrow's curriculum utility—that belongs to transition and horizon modeling.

3.5 What temporal model lies behind the teacher?

Narvekar and Stone explicitly formulate curriculum sequencing as an MDP, so its action value can in principle contain delayed transfer and training cost. Learning to Teach and El-Bouri also use discounted RL teachers over student transitions. PAIRED's environment designer participates in a multi-step construction and co-evolution process. But an RL algorithm alone does not establish that the learned policy captures meaningful long-term teaching effects; that depends on state, reward, horizon, and data.

Several recent LLM methods emphasize adaptive reactivity rather than an explicit multi-step teacher model. DUMP, SEC, and SPaCe update current arm utilities; Actor–Curator maximizes cumulative estimated actor improvements through a non-stationary bandit formulation. BMC explicitly models drift—including endogenous drift—in its learning-signal beliefs, but it still does not posit a full student-state transition model \(\widehat P^C(z'\mid z,a)\). Modeling endogenous non-stationarity is not the same as modeling full student dynamics. This lighter abstraction can be rational when outer transitions are scarce, state is hard to observe, and the action space is large.

\[\boxed{H^C\text{, the planning horizon, is a modeling choice.}}\]

4. Scaling pressures in LLM RL post-training

This section concerns curriculum mechanisms for RL-based LLM post-training, not curriculum learning for LLM pretraining in general. Pretraining also includes fixed, pacing-based, and interleaved curricula outside the pattern reviewed here.

Several recent RL post-training methods converge on lightweight bandit-style outer loops. I view the following as scaling pressures, not a claim that scale uniquely determines the algorithm.

The action space can be enormous. Candidate problems can number in the thousands or millions. Outer transitions can be unusually expensive. A decision may trigger multiple rollouts, verifier calls, a GRPO/PPO update, and evaluation, increasing \(\kappa^C\). Inner training exposes convenient signals. Reward, solve rate, advantage, and policy statistics already exist. Meaningful outer observations are few. A high-capacity teacher may not receive enough post-update transitions to justify its state and dynamics model.

\[\text{cheap observation}+\text{available proxy}+\text{bandit optimizer}\quad\text{becomes attractive}.\]

Clustering in SPaCe shares evidence; distribution arms in DUMP control complexity; SEC reuses advantage; BMC uses hierarchical problem relations to share evidence across an otherwise enormous prompt space; Actor–Curator adds neural generalization and a more target-aligned signal without building a full student-state MDP. These are different responses to the same outer-loop economics.

5. How to use the framework

\[\boxed{\mathcal E^C\longrightarrow\mathcal M^C}\]
\[\boxed{\begin{aligned}\textbf{State abstraction:}\quad&\mathcal Z^C\longrightarrow\widehat{\mathcal O}^C\\[1mm]\textbf{Action abstraction:}\quad&\mathcal A^C\longrightarrow\widehat{\mathcal A}^C\\[1mm]\textbf{Objective instrumentation:}\quad&J^C\longrightarrow\widehat R^C\\[1mm]\textbf{Dynamics abstraction:}\quad&P^C\longrightarrow(\widehat P^C,H^C)\end{aligned}}\]

The remaining pair, \((\Pi^C,\kappa^C)\), asks how the teacher learns under the cost imposed by those abstractions.

For a new curriculum method, I would ask seven questions:

  1. Observation: what does the teacher know about the student?
  2. Action: what training intervention can it make?
  3. Reward: what is treated as curriculum utility?
  4. Dynamics: how is the fact that teaching changes the learner represented?
  5. Optimizer: how are curriculum decisions updated?
  6. Horizon: is the teacher reactive or planning across future student states?
  7. Cost: how expensive is one useful teacher observation?

A non-stationary bandit is attractive when useful arms recur, student state is expensive to represent, immediate utility is informative enough, and outer transitions are scarce. A contextual teacher is attractive when a compact learner summary strongly predicts curriculum utility. A sequential RL or POMDP teacher becomes worth considering when prerequisites, transfer, interference, forgetting, or delayed effects make myopic utility systematically misleading—and when enough outer data exists to learn the extra structure.

A bandit teacher is not a less sophisticated point on a universal ladder toward RL; it is a different approximation to \(\mathcal E^C\). The right question is: which pieces of the teacher–student dynamics are worth paying to model?

Conclusion

Curriculum learning is not defined by easy-to-hard ordering. Once curriculum adapts, it becomes an outer decision system around a learner whose state is only partially visible and whose future changes in response to the teacher's own actions.

The universal questions are now compact: what does the teacher observe; what can it change; how is usefulness measured; how does the learner change; how is the teacher optimized; how far ahead does it reason; and what does this information cost?

The question is not whether curriculum learning should use a bandit or reinforcement learning. It is which parts of the teacher–student dynamics are worth modeling explicitly.

The central object is the value of a training intervention for the learner that exists now—and, when planning matters, for the learner that intervention will create next.

References

← Back to Artifacts & Blogs
研究笔记

谁来教
Learner?

一篇 framework-first review:当 curriculum 成为外层决策问题,teacher 在观察什么、控制什么、优化什么,又选择不显式建模什么?

郭道远2026 年 8 月约 17 分钟

Curriculum 不只是一串数据,而是对一个不断变化的 learner 施加的一连串干预。

假设 learner 可以学习 algebra 或 geometry。今天 geometry 带来的提升更大,即 \(\Delta J_t(\text{geometry})>\Delta J_t(\text{algebra})\);但 algebra 可能是 prerequisite,使明天学习 geometry 的价值大幅提高。于是

\[\text{最佳即时课程}\neq\text{最佳 curriculum 决策}.\]

这个小例子包含了主要困难。Curriculum 不等于 easy-to-hard 排序。一次课程会干预 learner,而 learner 的能力正因干预而变化;一旦选择规则根据反馈适应,curriculum design 就成为外层决策问题。

本文不是对 Curriculum Learning 全领域的综述。Fixed easy-to-hard schedule、self-paced learning、competence rule 与 pretraining-data curriculum 都是重要的相邻方向;这里聚焦被视为 adaptive outer decision process 的 curriculum mechanism。

在本文中,我用 adaptive curriculum optimization 作为描述性的 umbrella phrase。其中,RL for Curriculum 更窄地指 bandit、contextual decision 与 MDP/POMDP teacher:它们从反馈中学习 curriculum decision rule。全文围绕一个问题:teacher 观察什么、能改变什么、如何测量有用性,以及应该显式建模多少动态?

\[\boxed{\text{真实 learner dynamics}\rightarrow\text{部分观测与 proxy}\rightarrow\text{算法近似}\rightarrow\text{curriculum decision}}\]

1. Teacher 的决策问题

潜在 student state 与 teacher observation

令 \(z_t\) 表示 student 的潜在学习状态。它是概念对象,不一定等于参数向量 \(\theta_t\)。它可能包含技能掌握、遗忘知识、迁移准备度、optimizer history、任务干扰关系,以及哪些技能仍然可学。

Teacher 通常无法直接观察 \(z_t\),只能从训练历史构造 observation:

\[o_t^C\sim\mathsf O^C(\cdot\mid z_t,h_t),\qquad\text{或简写为}\qquad o_t^C=\Phi(h_t).\]

其中 \(h_t\) 可以包含近期 loss、evaluation、solve rate、advantage、probe performance、此前的 curriculum choice、gradient statistics 或 capability summary。计算得到的 learner summary 是对 student state 的观测,不会自动成为 Markov state。

Teacher action 与 student update

Curriculum action 是对 student 训练经验的干预:

\[a_t^C=\text{task、source、difficulty range、category、cluster、batch、problem、goal 或 environment}.\]

Teacher 通常不直接修改 student 参数,而是改变产生下一次 update 的数据、任务或环境:

\[\theta_{t+1}=\mathcal A\!\left(\theta_t,\mathcal B_t(a_t^C)\right).\]

Teacher transition、reward 与 objective

在概念层面,干预改变潜在 student state:

\[z_{t+1}\sim P^C(\cdot\mid z_t,a_t^C),\qquad a_t^C\rightarrow z_{t+1}\rightarrow\text{未来 curriculum utility}.\]

理想的 teacher signal 可以依赖干预前状态、action 与干预后状态,即 \(r_t^C=G(z_t,a_t^C,z_{t+1})\)。Teacher 可能优化最终 student quality,也可能优化累计 outer return:

\[\max_{\pi^C}\mathbb E[J(\theta_T)]\qquad\text{或}\qquad\max_{\pi^C}\mathbb E\!\left[\sum_{t=0}^{T-1}\gamma_C^t r_t^C\right].\]
\[\boxed{J^C\longrightarrow R^{C,*}\longrightarrow\widehat R^C}\]

这里,\(J^C\) 是最终训练目标;\(R^{C,*}\) 是与长期目标一致的局部 credit decomposition——如果这样的分解有意义;\(\widehat R^C\) 则是方法实际能测到的 signal。并非每个 curriculum objective 都有唯一合理的一步 teacher reward。实践中的 learning progress、validation gain、solve rate、advantage magnitude 或 policy-improvement estimate,都是对目标的 instrumentation。

latent student→teacher observation→intervention→student update→proxy signal

术语边界

\[\boxed{\begin{aligned}\text{Curriculum for RL}&=\text{curriculum mechanism}+\text{RL student},\\\text{RL for Curriculum}&=\text{feedback-learned outer decision rule}+\text{student learner}.\end{aligned}}\]

FastCuRL 使用分阶段的数据划分与 context extension 训练 reasoning model;其 outer schedule 是设计出来的,而非在线学习。DUMP、SEC、SPaCe 则根据训练反馈调整数据选择机制。Inner optimizer 是否为 GRPO 并不决定分类,outer mechanism 才决定。

Underlying process 与 method approximation

最干净的区分,是把我们理想上想推理的 reference teacher–student process,与具体方法实际实现的 abstraction 分开:

\[\boxed{\mathcal E^C=(\mathcal Z^C,\mathsf O^C,\mathcal A^C,P^C,J^C)}\]
\[\begin{aligned}\mathcal Z^C&:\text{ latent learner state},&\mathsf O^C&:\text{ observation kernel},&\mathcal A^C&:\text{ training interventions},\\P^C&:\text{ learner dynamics},&J^C&:\text{ ultimate training objective}.&&\end{aligned}\]

具体方法只实现对这一过程的近似:

\[\boxed{\mathcal M^C=(\widehat{\mathcal O}^C,\widehat{\mathcal A}^C,\widehat R^C,\widehat P^C,\Pi^C,H^C,\kappa^C)}\]
\[\begin{aligned}\widehat{\mathcal O}^C&:\text{ 实际使用的 student information},&\widehat{\mathcal A}^C&:\text{ 暴露给 teacher 的 action abstraction},\\\widehat R^C&:\text{ 可观测 curriculum signal},&\widehat P^C&:\text{ 假设或隐含的 learner dynamics},\\\Pi^C&:\text{ decision-update rule},&H^C&:\text{ effective planning horizon},&\kappa^C&:\text{ information 与 decision cost}.\end{aligned}\]

在本文中,我把每个 curriculum method 看作对共同 reference process \(\mathcal E^C\) 的近似,并由 method-specific abstraction \(\mathcal M^C\) 实现。我们理想上想推理的学习过程,与算法实际建模的信息和结构之间的区分,才是这篇 review 的组织工具。

2. 为什么 Curriculum 是困难的 Outer-loop Learning Problem?

下面六个困难可以归入三组:dynamics、information 与 statistics/systems。

\[\boxed{\begin{array}{lll}\textbf{Dynamics}&\text{sequentiality}&\text{endogenous non-stationarity}\\[2mm]\textbf{Information}&\text{partial observability}&\text{proxy reward}\\[2mm]\textbf{Statistics / systems}&\text{expensive feedback}&\text{large action spaces}\end{array}}\]

2.1 Sequentiality

即时有用性 \(\mathbb E[r_t^C\mid z_t,a_t^C]\) 与长期 curriculum value \(Q^C(z_t,a_t^C)\) 可能不一致。Algebra 能为 geometry 建立 prerequisite;简单样本可能形成很久以后才使用的表征;回访旧任务可以抵抗 forgetting;一个 domain 也可能向另一个 domain 迁移。

\[a_t^C\rightarrow z_{t+1}\rightarrow z_{t+2}\rightarrow\cdots\]

当今天的 lesson 改变明天可以高效学习什么时,curriculum 才真正成为 sequential problem。

2.2 Endogenous non-stationarity

Non-stationary bandit 可以追踪 \(\mu_k(t)=\mathbb E[r_t^C\mid a_t^C=k]\)。但 curriculum 中的 drift 往往是内生的:

\[a_t^C\rightarrow z_{t+1}\rightarrow\mu_k(t+1).\]

Teacher 不只是追踪变化的 environment;它过去的干预也创造了后来需要适应的变化。Non-stationary bandit 可以追踪移动的 arm utility,却不一定显式建模今天的 action 如何因果地改变明天的 arms。这是一种建模简化,并非缺陷。

\[\boxed{\text{tracking }\mu_k(t)\neq\text{modeling how }a_t^C\text{ changes }\mu_k(t+1)}\]

Curriculum non-stationarity 往往是 endogenous:teacher 改变了产生未来 arm value 的 learner。Utility-tracking approximation 正是把这层因果关系留在模型之外。

近期 LLM curriculum 工作已经开始显式采用这一视角。Manifold Bandits 将 problem sampling 建模为具有 endogenous non-stationarity 的 structured bandit:被选中的 prompts 会影响 learning signal 随 policy latent task geometry 的后续演化。

2.3 Partial observability

Teacher 看到的是 \(o_t^C=\Phi(h_t)\),而不是完整学习状态。两个近期 loss 相同的 student,可能在 mastery、transfer potential、forgotten capabilities 或 gradient geometry 上完全不同。当现有 learner summary 不足以预测 curriculum consequence 时,POMDP 是一种自然的 reference abstraction。实践问题不是“有没有 state vector”,而是“这个 student representation 对 curriculum decision 是否足够有信息”。

2.4 Proxy reward 与 myopia

需要区分两个问题。Proxy alignment 问 \(\widehat r_t^C\) 是否预测理想的最终表现 \(J(\theta_T)\);planning 问即使 immediate gain 正确,贪心最大化它是否给出最佳 sequence。即使

\[\widehat r_t^C=J(\theta_{t+1})-J(\theta_t),\]

prerequisite、transfer 与 delayed effect 仍然存在。对齐的一步估计不会自动消除 planning problem。

2.5 昂贵的 outer-loop feedback

select data→train student→roll out / evaluate→measure utility

一次 curriculum transition 可能包含大量 inner training samples,特别是 LLM post-training 中常有 \(N_{\text{outer decisions}}\ll N_{\text{inner samples}}\)。丰富的 outer model 可能表达力很强,却无法从现有 teacher transitions 中估计出来。Teacher sample efficiency 本身就是问题定义的一部分。

2.6 大型且结构化的 action space

Teacher 可能在十几个 tasks 中选择,也可能面对数百万 problems。大型 action space 提高探索成本并造成 sparse feedback,但 action 之间也有结构:category、cluster、embedding、hierarchy 与 generative parameterization 可以在相似干预之间共享证据。Teacher 选择行动粒度,本身就是统计建模决定。

因此,文献可以被理解为从 \(\mathcal E^C\) 到 \(\mathcal M^C\) 的不同近似:有些压缩 student information,有些简化 intervention space、proxy、transition structure、horizon 或 cost。这是本文的组织框架,并非领域内标准 notation。

3. 文献如何建模 Teacher?

表格先给出地图;随后正文解释其中反复出现的 modeling choices。

Method年代 / domainTeacher observationCurriculum actionTeacher signalOuter optimizerOuter temporal model
Graves et al.2017, supervisedTask reward history;无显式 student stateTask / data sourcePrediction 或 complexity gainExp3.SBandit;无 transition model
Matiisen et al.2017, supervised / RLTask performance slopeSubtaskLearning progress,含 forgettingBandit-style teacherReactive progress tracking
Learning to Teach2018, supervisedStudent、data、history features保留 / 丢弃 example达到目标 accuracy 的时间REINFORCEDiscounted teaching episode
Narvekar & Stone2019, RL transferAgent knowledge representationSource task / transferNegative time-to-learnSARSA(\(\lambda\))Explicit curriculum MDP
El-Bouri et al.2020, supervisedEncoded student weightsBatch / continuous data regionValidation-weighted progressDQN / DDPGDiscounted teaching episode
PAIRED2020, environment designPartially built environmentConstruct environmentProtagonist–antagonist return gapPPOSequential environment construction
FastCuRL2025, LLM RFT — boundary caseFixed stage indexData segment / context stageDesigned scheduleFixed schedulerPredefined stage sequence
DUMP2025, LLM RFT近期 distribution advantage 与 countDistribution mixtureMean absolute advantageUCBNon-stationary arm utilities
SEC2025, LLM RFTCategory valuesReasoning categoryMean absolute advantageBoltzmann + TD(0)Bootstrapped category utilities
SPaCe2025/26, LLM RFTCluster solve rate、pull、stagnationData clusterCorrectness / hardness / progressThompson SamplingNon-stationary cluster utilities
Manifold Bandits / BMC2026, LLM RFTPrompt-level learning-signal belief + latent task treeHierarchically related promptsRollout-reward variabilityHierarchical Thompson SamplingStructured bandit + endogenous drift
Actor–Curator2026, LLM RFTCandidate problems;online utilityProblem probabilitiesPer-problem policy-improvement estimateNeural OSMD + clippingBandit over one-update improvement

3.1 Teacher 对 Student 知道什么?

早期工作大致探索三种策略:不使用显式 state、手工设计 behavioral summary,以及学习 state embedding。Graves 只通过 task reward history 表示 student change;Kumar 用 prototype sentences 上的 likelihood 构造紧凑行为摘要;Reinforcement Teaching 从 probe behavior 学习 embedding。其他 state-conditioned teacher 还会使用 student、data、history、knowledge 或 parameter features,具体差异交给表格索引。

近期 LLM scheduler 往往不尝试学习整个 student 的紧凑 Markov representation。Learner 的变化不再通过显式状态表征,而是反映在随训练不断变化的 distribution、category、cluster 或 problem utility estimate 中。DUMP 的 advantage window 与 SPaCe 的 solve-rate statistics 是这一路线的代表。

\[\boxed{\text{student-state conditioning}\qquad\text{与}\qquad\text{utility tracking}}\]

这是两种常见策略,而不是同一条连续轴的两端;方法可以一边 condition on student features,一边追踪 non-stationary utility。BMC 还展示了第三种选择:不显式表示 student state,却表示 actions 之间的结构。

3.2 Teacher 能选择什么?

Pre-LLM 工作的 action 包含 task/source(Graves、Matiisen)、data/noise bin(Kumar)、batch 或 difficulty-sorted data 的连续区域(El-Bouri)、source task 与 transfer choice(Narvekar、Stone),以及构造 environment(PAIRED)。LLM scale 下,DUMP 作用于 distribution mixture,SEC 选择 reasoning category,SPaCe 选择 semantic/difficulty cluster,Actor–Curator 则给 individual candidate problems 分配选择概率。

\[\text{coarse shared action}\quad\longleftrightarrow\quad\text{fine local action}\]
\[\text{independent actions}\quad\longleftrightarrow\quad\text{structured、related actions}\]

Granularity 与 structure 是两条不同的轴。Coarse action 天然共享证据;fine action 提供精确干预,却需要更多反馈,除非方法能利用 actions 之间的关系。Actor–Curator 在 problem representation 之间泛化;BMC 则根据 policy embedding 构造 hierarchical task tree,使证据能在相关 prompts 之间传播。因此 \(\widehat{\mathcal A}^C\) 同时规定 intervention resolution,以及 teacher 能利用怎样的 cross-action structure。

3.3 什么算一堂有用的课?

Learning progress。 Graves 等从 prediction gain 或 model-complexity change 构造 reward;Matiisen 等使用 task-performance slope,使 performance 下降可以触发复习;Reinforcement Teaching 也用 learning-progress shaping。它们试图定位 learner 不断移动的能力边界:

\[\widehat r_t^C\approx m_{t+1}-m_t.\]

Training efficiency。 Learning to Teach 的 data-teaching experiment 使用 terminal time-to-threshold signal;Narvekar 与 Stone 把 negative time-to-learn 当作 curriculum cost。Lesson 的价值体现在减少训练成本。

Regret-like utility。 PAIRED 用 protagonist 与 antagonist 的 return gap 近似 regret,训练 environment designer。它寻找能暴露弱点、但仍可被另一 agent 解决的 environment;这是特定的 unsupervised environment-design construction,不应泛化为所有 curriculum reward。

LLM scale 的 inner-learning signal。 DUMP 与 SEC 用 expected absolute advantage 近似 immediate learnability:

\[\widehat r_t^C\approx\mathbb E[|\widehat A|].\]

它表示现有 RL training 中即时 policy-learning signal 的存在,不直接度量长期 curriculum value。SPaCe 则把 binary correctness 聚合为 cluster solve rate,转成 hardness,并降低长期 stagnant clusters 的选择权重。

Policy-improvement estimate。 Actor–Curator 根据更新前后 policy 估计每道 problem 对 \(J(\pi_{t+1})-J(\pi_t)\) 的贡献。它是一次 actor update 所诱发 performance gain 的 first-order credit-assignment signal,并不会自动解决 long-horizon planning。

\[\boxed{\text{signal semantics}\quad\times\quad\text{measurement cost}}\]

Learning progress、training efficiency、regret、inner-optimization signal 与 policy improvement 描述的是不同 semantics,而不是单调 progression。每种 signal 都有自己的 cost、variance、alignment 与 horizon;昂贵的 measurement 不会自动更 aligned。BMC 还提出一个有用的诊断问题:instrument 测量的是 productivity(learning signal)、diversity(task-manifold coverage),还是 utility(evaluation relevance)?这三种 curriculum value 并不可互换。

3.4 Teacher 如何学习并适应?

Fixed scheduler 预先指定 \(\pi_t^C\),例如 FastCuRL 的 staged segmentation 与 context extension。Bandit teacher 则学习 action utility:Graves 使用 non-stationary Exp3.S;DUMP 用带 recent advantage window 的 UCB;SEC 用 Boltzmann sampling 与 TD(0) category-value update;SPaCe 用 progress-aware Thompson Sampling;Actor–Curator 使用 neural function approximation 与源自 OSMD、带 PPO-style clipping 的 curator objective。BMC 在 latent task tree 上运行 hierarchical Thompson Sampling,并结合 Bayesian filtering 与 bottom-up empirical-Bayes propagation,让相关 prompts 共享证据。

Contextual/state-conditioned teacher 依据 \(a_t^C\sim\pi^C(\cdot\mid o_t^C)\) 决策。Learning to Teach 在 student/data/history features 上使用 REINFORCE;Kumar 在 prototype-sentence likelihood 上使用 DQN;El-Bouri 用 DDPG 控制连续 batch location/width,并比较 DQN discrete variant;Narvekar 与 Stone 用 SARSA(\(\lambda\)) 学 curriculum-MDP policy;PAIRED 用 PPO 训练 environment-design agents;Reinforcement Teaching 使用 DQN/Double DQN teacher 与 learned behavioral embedding。

这些算法尝试利用 state dependence,有时也利用 delayed effect。但使用 RL optimizer 并不保证 observation 是 Markov,也不保证真正学到了有用的长期规划。

Explicit-state adaptation 因 \(o_t^C\) 改变而换 action:\(a_t^C\sim\pi^C(\cdot\mid o_t^C)\)。Implicit utility tracking 因 \(\widehat\mu_k(t)\) 移动而换 action。Graves 在 non-stationary bandit 中共享权重;DUMP 使用 sliding window;SEC 用 TD(0) bootstrap category values;SPaCe 更新 solve-rate 与 stagnation statistics;Actor–Curator 在 actor update 后接收 online partial-feedback estimate。

显式 student-state modeling 与 non-stationary arm tracking,是适应 changing learner 的两种不同方式。二者都不会自动建模“今天 action 如何因果地改变明天 curriculum utility”;后者属于 transition 与 horizon modeling。

3.5 Teacher 背后是什么 temporal model?

Narvekar 与 Stone 显式把 curriculum sequencing 写成 MDP,因此 action value 原则上可以包含 delayed transfer 与 training cost。Learning to Teach 和 El-Bouri 也在 student transitions 上使用 discounted RL teacher。PAIRED 的 environment designer 参与 multi-step construction 与 co-evolution。但仅仅使用 RL algorithm 不能证明学到的 policy 捕捉了长期 teaching effects;它仍取决于 state、reward、horizon 与数据。

若干近期 LLM 方法强调 adaptive reactivity,而非显式 multi-step teacher model。DUMP、SEC、SPaCe 更新当前 arm utility;Actor–Curator 通过 non-stationary bandit 最大化累计 estimated actor improvement。BMC 显式追踪 learning-signal belief 的 drift,包括 endogenous drift,但仍未假设完整的 student-state transition model \(\widehat P^C(z'\mid z,a)\)。建模 endogenous non-stationarity,不等于建模完整 student dynamics。当 outer transition 稀缺、state 难观测、action space 巨大时,这种较轻的 abstraction 可能更合理。

\[\boxed{H^C\text{(规划时域)也是建模选择。}}\]

4. LLM RL Post-training 中的 Scaling Pressures

本节讨论的是 RL-based LLM post-training 中的 curriculum mechanism,而不是整个 LLM pretraining curriculum。后者还包含 fixed、pacing-based 与 interleaved curricula,不属于这里概括的模式。

若干近期 RL post-training 方法不约而同地选择轻量 bandit-style outer loop。我把下面这些理解为 scaling pressures,而不是“scale 唯一决定算法”的历史断言。

Action space 可能极其庞大:候选 problems 可达成千上万乃至更多。Outer transition 可能格外昂贵:一次决策可能触发多次 rollout、verifier call、GRPO/PPO update 与 evaluation,从而提高 \(\kappa^C\)。Inner training 暴露方便的 signal:reward、solve rate、advantage 与 policy statistics 已经存在。有意义的 outer observation 较少:高容量 teacher 可能没有足够 post-update transitions 来支撑自己的 state/dynamics model。

\[\text{cheap observation}+\text{available proxy}+\text{bandit optimizer}\quad\text{因此很有吸引力}.\]

SPaCe 用 clustering 共享证据;DUMP 用 distribution arms 控制复杂度;SEC 复用 advantage;BMC 借助 problems 之间的 hierarchy,在原本极大的 prompt space 中传播证据;Actor–Curator 加入 neural generalization 与更贴近目标的 signal,却不构建完整 student-state MDP。这些是对相同 outer-loop economics 的不同回应。

5. 如何使用这个 Framework

\[\boxed{\mathcal E^C\longrightarrow\mathcal M^C}\]
\[\boxed{\begin{aligned}\textbf{State abstraction:}\quad&\mathcal Z^C\longrightarrow\widehat{\mathcal O}^C\\[1mm]\textbf{Action abstraction:}\quad&\mathcal A^C\longrightarrow\widehat{\mathcal A}^C\\[1mm]\textbf{Objective instrumentation:}\quad&J^C\longrightarrow\widehat R^C\\[1mm]\textbf{Dynamics abstraction:}\quad&P^C\longrightarrow(\widehat P^C,H^C)\end{aligned}}\]

剩下的 \((\Pi^C,\kappa^C)\) 回答另一个问题:在这些 abstraction 与可用成本之下,teacher 如何学习?

面对新的 curriculum method,我会问七个问题:

  1. Observation:teacher 对 student 知道什么?
  2. Action:它能施加什么训练干预?
  3. Reward:什么 signal 被当作 curriculum utility?
  4. Dynamics:方法如何表示 teaching 会改变 learner?
  5. Optimizer:curriculum decision 如何更新?
  6. Horizon:teacher 是即时反应,还是跨未来 student states 规划?
  7. Cost:获得一次有用 teacher observation 有多贵?

当有用 arms 会重复出现、student state 难以表征、immediate utility 足够有信息、outer transition 稀缺时,non-stationary bandit 很有吸引力。当紧凑 learner summary 能强力预测 curriculum utility 时,contextual teacher 值得采用。当 prerequisite、transfer、interference、forgetting 或 delayed effect 让 myopic utility 系统性误导,而且 outer data 足够学习额外结构时,sequential RL/POMDP teacher 才开始值得。

Bandit teacher 不是通往 RL 的“低配阶段”,而是对 \(\mathcal E^C\) 的另一种近似。正确问题是:teacher–student dynamics 中,哪些部分值得付出成本去显式建模?

结论

Curriculum Learning 并不由 easy-to-hard ordering 定义。一旦 curriculum 能够适应,它就成为围绕 learner 的 outer decision system:learner state 只有部分可见,而它的未来又会响应 teacher 自己的 action。

通用问题由此变得清楚:teacher 观察什么;能改变什么;如何测量有用性;干预后 learner 怎样变化;teacher 如何优化;它向前推理多远;这些信息又需要多少成本?

问题不是 curriculum learning 应该使用 bandit 还是 reinforcement learning,而是 teacher–student dynamics 中哪些部分值得被显式建模。

核心对象是一次 training intervention 对“此刻这个 learner”的价值;当 planning 重要时,还包括它对这次干预将创造出的“下一个 learner”的价值。

参考资料

← 返回 Artifacts & Blogs