Project reflection

Beyond the Algorithm Label

How an apparently successful RL experiment taught me to stop treating algorithm names as causal explanations.

Daoyuan GuoAugust 20269 min read

This essay is a personal reflection on a collaborative project with Zhouan Shen and Yilin Liu. It draws on our report and the project repository, but focuses on the questions I would ask after the project rather than presenting the report again.

R@1: 0.4006 → 0.4434. At first, this looked like our clearest result. It eventually became the result I trusted the least.

We called the winning configuration “UCB-inspired.” My first reading was simple: a more structured exploration rule had beaten greedy search. Then I traced the implementation and found that the run had changed two things at once. It added a count-based bonus, but it also removed the original Ornstein–Uhlenbeck (OU) noise.

After that, “which algorithm won?” no longer felt like the right question. I wanted to know what the experiment had actually isolated.

One limitation applies to every number below: these were single-run comparisons. I now treat them as clues about mechanisms, not estimates of method superiority. The experiments were good enough to raise questions, not good enough to settle them.

1. What had actually changed?

Table2Charts generates a chart as a grammar-constrained token sequence. A partial sequence is a prefix; the model scores every legal next token; search selects a continuation and retains alternatives; visited prefixes enter replay memory; the scorer is then updated.

prefix→legal actions→search→replay→training
exploration  → visited states
replay       → training distribution
targets      → supervision semantics
search       → reachable outputs
metric       → definition of success

This is sequential, but not textbook online RL. Transitions are deterministic, legal actions come from a grammar, positive targets come from corpus charts, and there is no conventional return backup from user interaction. The problem was not calling it RL-style. The problem was allowing RL vocabulary to do explanatory work that the implementation had not earned.

2. Exploration was also data collection

I initially thought epsilon exploration changed action selection. It did, but that was only its immediate effect. Because selected prefixes were later stored in replay memory, the policy also changed the dataset used for subsequent training:

\[\pi\longrightarrow d^{\pi}\longrightarrow D_{\pi}\longrightarrow\pi'\]

What we observed: hard-target greedy obtained R@1=0.4006, while ε=0.20 obtained 0.4253. The effect was non-monotonic: ε=0.10 was close to baseline and ε=0.30 was weaker than ε=0.20.

One plausible interpretation: moderate top-M exploration may have exposed training to prefixes closer to those encountered under model-driven search, potentially reducing the mismatch between supervised and inference-time prefixes.

What we did not identify: we did not log replay composition, prefix-depth coverage, or recovery after an incorrect branch. The data-distribution explanation is therefore a hypothesis, not a demonstrated mechanism. A broader lesson remains: in sequential ML systems, changing the decision rule can also change the data on which the next model is trained.

3. Our “UCB” experiment did not isolate UCB

When I first saw R@1=0.4434, I treated the method name as the explanation. Code inspection forced a revision. Classical UCB attaches optimism to uncertainty estimates that become informative as evidence accumulates for comparable decisions. A raw state–action count in a sparsely revisited prefix tree does not obviously inherit that calibrated uncertainty interpretation.

At the same time, the run removed OU noise. OU perturbations have a natural role in continuous control; here they were added to scores for discrete tokens whose valid set changes by prefix. There is no obvious geometry in which temporal smoothness across token coordinates should be useful.

Observed: the combined configuration scored highest in one run. Identified: the experiment compared two bundles, not UCB versus no UCB. Unresolved: whether the gain came from the bonus, removal of noise, their interaction, or run-to-run variation.

The highest score was real. The causal story attached to it was not identified.

The minimum follow-up is a 2×2 factorial ablation—OU/no OU crossed with bonus/no bonus—repeated across seeds. Until then, “UCB won” is a label, not an explanation.

4. “Dense reward” chose which errors should look similar

I thought we were improving reward shaping. At code level, we were mostly changing labels. The targets were 0.95 for the exact next action, 0.35 for a field of the same type, 0.15 for another field, 0.10 for the same token type, and 0.05 otherwise.

This replaces one-hot supervision with a hand-designed action geometry—an implicit similarity kernel \(K(a,a^*)\). Once the target became soft, we were no longer only asking which action was correct. We were also deciding which wrong actions should count as similar. Why should “same field type” be closer than “same token type”? Does that reflect grammar, meaning, or user utility? We had not tested the answer.

SettingR@1R@3R@5R@10Unique rate
Hard + greedy0.40060.75980.83120.90940.0100
Conservative soft + greedy0.41570.74350.83100.90050.0131
Aggressive soft + greedy0.40780.73640.80280.87110.0174

Single-run results on the regenerated Plotly setting; patterns, not confidence-qualified estimates.

Soft targets sometimes improved R@1 and increased diversity while weakening broader top-k recovery. They may have reduced the ranking margin between the logged action and related alternatives, but we did not establish whether those alternatives were useful recommendations; exact recall cannot answer that.

5. The metric was quietly choosing the winner

Complete chart-sequence recall asks whether a chart recorded in the corpus appears in the top-k generated sequences. It is reproducible and useful for evaluating corpus reconstruction. It does not directly measure whether a recommendation communicates the data well, helps a user discover an insight, or matches an unstated analytic goal.

We had started by asking which algorithm produced the best recommendations. By this point, I was no longer sure exact sequence recall alone was enough to support claims about recommendation quality. A benchmark can quietly turn an open-ended recommendation problem into a reconstruction problem: maximizing R@k means maximizing the probability of recovering a logged chart, not maximizing expected human utility.

Hard targets, a prefix-consistency critic, and exact recall are aligned with one another. Dense targets that admit related actions may raise diversity while losing exact recall. The metric did not merely observe the winner; it helped choose which behavior could win.

A stronger evaluation needs three distinct axes:

  1. Recovery: exact R@k for comparability with Table2Charts.
  2. Validity and diversity: grammar validity, semantic compatibility, uniqueness, and chart-type balance.
  3. Utility: blinded human preference or success on a concrete analytic task.

6. Architecture names did not determine the learning algorithm

I also thought adding an actor made the system more actor–critic-like. But the policy head was not trained from a conventional TD advantage over trajectory returns; it learned from prefix-action supervision and the critic’s current signal. Actor-only ranking was weak, and blending it with the critic did not improve global ranking.

This does not identify a failure of actor–critic for visualization. It identifies a naming mistake in the other direction: a policy-shaped output head does not become actor–critic merely because we call one output “actor” and another “critic.” The objective and credit-assignment mechanism matter more than the architecture label.

7. What I would measure before naming the next algorithm

Did it work? Why did it work? Was that the right thing to optimize?

Identification — did it work?

  • Factorial ablations that change one mechanism at a time.
  • Multiple seeds, uncertainty intervals, and matched search budgets.

Mechanism — why did it work?

  • Replay composition, prefix-depth coverage, and recovery after mistakes.
  • Interactions among exploration, beam size, frontier size, and expansion limit.

Objective — was it the right thing to optimize?

  • Exact recovery, semantic validity, and diversity reported separately.
  • Human preference or task utility for claims about recommendation quality.

Together, these measurements would let me test the chain shown at the start instead of treating it as a finished explanation.

8. The rule I am taking forward

When a result improves, first list everything that changed.

If that list contains more than the mechanism named in the experiment, I no longer treat the method name as the explanation. The score was a measurement. The label was a name. The causal story still had to be earned.

Project collaborators: Zhouan Shen and Yilin Liu. This retrospective represents my own interpretation of the collaborative work. Results are drawn from the project report and repository and are subject to its documented experimental limitations.
← Back to homepage
项目反思

算法标签之外

一次看似成功的 RL 实验,如何让我不再把算法名称直接当作因果解释。

郭道远2026 年 8 月约 9 分钟

本文是我对一次合作项目的个人反思。项目由我与 Shen Zhouan、Liu Yilin 共同完成。文章依据我们的报告和项目仓库,但重点不是复述报告,而是追问项目结束后仍值得继续分析的问题。

R@1:0.4006 → 0.4434。起初,它看起来是项目中最清楚的结果;后来,它却成了我最不信任的结果。

我们把最佳配置称为 “UCB-inspired”。我最初的理解很简单:更有结构的 exploration 胜过了 greedy search。后来沿着代码检查,我才发现这个实验同时改变了两件事:增加 count-based bonus,也移除了原来的 Ornstein–Uhlenbeck(OU)noise。

从那以后,“哪个算法赢了”就不再是我最想问的问题。我更想知道,这个实验到底隔离出了什么。

下文所有数字都有同一个限制:这些是 single-run comparisons。我现在把它们视为寻找机制的线索,而不是方法优劣的稳定估计。实验足以提出问题,却不足以终结问题。

1. 实际上改变了什么?

Table2Charts 将图表生成为受 grammar 约束的 token sequence。模型为当前 prefix 的每个合法 next token 评分;search 选择 continuation 并保留其他候选;访问过的 prefixes 进入 replay memory;随后 scorer 被更新。

prefix→合法动作→search→replay→训练
exploration  → 访问的状态
replay       → 训练分布
targets      → 监督语义
search       → 可到达的输出
metric       → 成功的定义

这个过程具有序列决策形式,却不是典型的 online RL:转移是确定性的,合法动作来自 grammar,正样本来自语料中的图表,也没有由用户交互产生的标准 return backup。问题不在于称它为 RL-style,而在于让 RL 术语替尚未被实现证明的理论机制完成解释工作。

2. Exploration 同时也是 data collection

我最初以为 epsilon exploration 改变的是 action selection。的确如此,但那只是即时作用。因为选中的 prefix 随后会进入 replay memory,policy 也改变了下一轮训练使用的数据:

\[\pi\longrightarrow d^{\pi}\longrightarrow D_{\pi}\longrightarrow\pi'\]

我们观察到:hard-target greedy 的 R@1=0.4006,ε=0.20 时为 0.4253;效果并不单调,ε=0.10 接近基线,ε=0.30 又弱于 ε=0.20。

一种合理解释:适度的 top-M exploration 可能让训练接触到更接近模型自主搜索时遇到的 prefixes,从而可能缩小监督数据与推理数据之间的分布偏差。

我们尚未识别:项目没有记录 replay memory 的构成、不同 prefix 深度的覆盖率,也没有测量走错分支后的恢复能力。因此,“改变训练分布”仍是假设,不是已证实的机制。能够确定的普遍问题是:在序列式机器学习系统中,改变决策规则,也可能同时改变下一轮模型面对的训练分布。

3. “UCB” 实验并没有隔离 UCB

第一次看到 R@1=0.4434 时,我把方法名称当成了因果解释。代码检查迫使我修改判断。经典 UCB 将 optimism 建立在 uncertainty estimate 上,而这种估计需要随着可比较决策的证据积累才变得有意义。对于很少重复访问的 prefix tree,原始 state–action count 并不会自然获得这种经过校准的不确定性解释。

与此同时,该 run 移除了 OU noise。OU perturbation 在 continuous control 中有自然用途;但这里的动作是离散 token,合法集合随 prefix 变化,noise 则直接叠加到 action score。不同 token 坐标之间不存在明显的连续几何关系。

观察结果:组合配置在一次运行中得分最高。已经识别:实验比较的是两组捆绑改动,而不是 UCB 与 no UCB。尚未解决:提升来自 bonus、移除 noise、两者交互,还是不同运行之间的随机波动。

最高分是真实的;附着在这个数字上的因果故事并没有被识别。

最基本的后续实验应采用 2×2 factorial ablation:OU / no OU 与 bonus / no bonus 正交组合,并跨多个 seeds 重复。在此之前,“UCB 赢了”只是标签,不是解释。

4. “Dense reward” 在规定哪些错误彼此相似

我原以为我们在改进 reward shaping;从代码看,我们主要改变了 labels:exact next action 为 0.95,相同字段类型为 0.35,其他字段为 0.15,相同 token 类型为 0.10,其余为 0.05。

这相当于把 one-hot supervision 换成一套人工设计的 action geometry——隐含的 similarity kernel \(K(a,a^*)\)。当 target 变软之后,我们不只是在问哪个动作正确,也是在决定哪些错误动作应该被视为相似。为什么 same field type 应比 same token type 更接近?这反映的是 grammar、语义,还是用户效用?我们没有验证答案。

SettingR@1R@3R@5R@10Unique rate
Hard + greedy0.40060.75980.83120.90940.0100
Conservative soft + greedy0.41570.74350.83100.90050.0131
Aggressive soft + greedy0.40780.73640.80280.87110.0174

Regenerated Plotly setting 的 single-run results;它们是 pattern,而非带置信区间的稳定估计。

Soft targets 有时提升 R@1 和多样性,却削弱更广的 top-k 恢复率。它可能缩小了已记录动作与相关替代动作之间的排序间隔,但我们没有识别这些替代动作是否更有用;exact recall 无法回答这一点。

5. 指标在悄悄选择赢家

Complete chart-sequence recall 检查语料中记录的图表是否出现在 top-k 生成序列中。它可复现,也适合评价语料重建;但它不直接衡量图表是否良好表达数据、是否帮助用户获得洞察,或是否符合一个未被记录的分析目标。

项目开始时,我们问哪种算法能产生更好的推荐。到这里,我已经不再确定 exact sequence recall 单独一项是否足以支持关于推荐质量的结论。Benchmark 可以在没有明确说明的情况下,把开放式推荐问题变成 reconstruction problem:最大化 R@k 是最大化恢复已记录图表的概率,并不等于最大化预期的用户效用。

Hard target、prefix-consistency critic 与 exact recall 彼此对齐;允许相关动作的 dense target 可能提高多样性,却损失 exact recall。指标并不只是观察谁赢了;它也参与决定什么样的行为有资格被称为“赢”。

更完整的评估应分为三个独立轴:

  1. Recovery:保留 exact R@k,以便与 Table2Charts 比较。
  2. Validity 与 diversity:grammar validity、semantic compatibility、uniqueness 和 chart-type balance。
  3. Utility:blind human preference 或具体分析任务上的成功率。

6. 架构名称并不决定学习算法

我也曾认为增加 actor 会让系统更接近 actor–critic。但 policy head 并非通过 trajectory return 上的标准 TD advantage 训练,而是从 prefix-action supervision 与 critic 当前信号学习。Actor-only ranking 很弱,混合 actor 与 critic 也没有改善全局排序。

这并未证明 actor–critic 不适合可视化;它揭示的是另一种命名错误:仅仅把两个输出分别称为 actor 与 critic,并不会让 policy-shaped output head 自动成为 actor–critic。学习目标与 credit assignment 机制比架构标签更重要。

7. 在命名下一个算法前,我会先测量什么

它有效吗?为什么有效?它优化的真是正确目标吗?

效果识别——它有效吗?

  • 每次只改变一个机制的 factorial ablations。
  • 多个 seeds、uncertainty intervals 与一致的 search budgets。

机制分析——为什么有效?

  • Replay memory 的构成、prefix 深度覆盖率与出错后的恢复能力。
  • Exploration 和 beam size、frontier size、expansion limit 的交互。

目标审查——它优化的真是正确目标吗?

  • 分别报告 exact recovery、semantic validity 与多样性。
  • 若要声称推荐质量提升,则加入人类偏好或任务效用。

有了这些测量,我才能真正检验文章开头的因果链,而不是把它直接写成最终解释。

8. 我会带进下一项研究的规则

当一个结果变好时,先列出所有发生变化的东西。

如果这份清单不只包含实验名称所指的那个机制,我就不再把方法名称直接当成解释。分数是测量结果,标签只是名称,而因果故事仍需由证据赢得。

项目合作者:Shen Zhouan、Liu Yilin。本文仅代表我对合作项目的个人理解。文中结果来自项目报告与仓库,并受到报告中所述实验限制的约束。
← 返回主页