Research note

Two Measurements,
One Learner

Rewards measure an option against a numerical scale. Preferences measure options against one another. If a learner receives both, how can it use them without making either kind of information disappear?

Daoyuan Guo March 2026 15 min read

One evaluator gives answer A an 8 out of 10. Another, shown A and B together, chooses B. What exactly have these observations told us?

Both are judgments about “goodness,” but they are not interchangeable. The score locates A on an evaluator’s numerical scale. The comparison locates B relative to A in a particular choice. It does not tell us how far apart they are, and the score does not tell us whether the second evaluator would use the scale in the same way.

This small puzzle is the reason I find “using rewards and preferences together” harder than it first sounds. If a learning system already receives both, what should it mean to fuse them? My working thesis is that fusion should not merely convert every signal into the same label format. The more demanding goal is to preserve what each measurement tells us while allowing all evidence to influence one coherent learning and control process.

1. Two measurements of “good”

Start with a stochastic bandit and assign arm \(i\) a latent value \(u_i\). A numerical reward is an outcome associated with one arm:

\[Y_i\sim P_R(\,\cdot\mid u_i),\qquad\text{for example }Y_i=u_i+\varepsilon_i.\]

A pairwise label is instead a relational outcome associated with two arms:

\[Z_{ij}\sim\operatorname{Bernoulli}\!\left(\mu_P(u_i-u_j)\right).\]

The binary observation \(Z_{ij}\) does not reveal a numerical utility difference. Only after choosing a link \(\mu_P\) and collecting enough comparisons can that difference be inferred. Reward data expose numerical magnitude under their observation model; preference data directly expose a noisy ordering event.

There is another distinction hiding here. The raw statement \(i\succ j\) constrains an ordering. Writing \(\Pr(i\succ j)=\sigma(\beta(u_i-u_j))\) adds a quantitative model of how latent differences become choices. If \(\beta\) is unknown, rescaling the utilities and inversely rescaling \(\beta\) leaves the probabilities unchanged. The observation format did not give us that scale; the parametric model did.

The invariance makes the distinction sharper. In the usual difference model, replacing every utility by \(u_i+c\) leaves all preference probabilities unchanged:

\[(u_i+c)-(u_j+c)=u_i-u_j.\]

Depending on \(P_R\), rewards can retain level information that comparisons remove. A reward anchors an alternative to a scale; a preference anchors one alternative to another.

With \(u_i=x_i^\top\phi\), this becomes geometric:

\[Y_i\sim P_R(\cdot\mid x_i^\top\phi),\qquad Z_{ij}\sim P_P\!\left(\cdot\mid(x_i-x_j)^\top\phi\right).\]

The reward channel measures along \(x_i\); the preference channel measures along \(x_i-x_j\). They are not two noisy copies of one observation. They are different observation operators acting on what may—or may not—be a shared latent object.

2. What should fusion mean?

I find it useful to distinguish four increasingly demanding meanings. This is my organization of the problem, not established terminology.

Format fusion

Convert heterogeneous observations into one representation. For example, \(r_A>r_B\Rightarrow A\succ B\). This is convenient, and sometimes appropriate, but \(10>9\) and \(100>-100\) both become the same binary order. Magnitude may be compressed.

Statistical fusion

Keep distinct observation models while letting them inform a shared latent object:

\[\ell(\phi)=\sum_{y\in D_R}\ell_R(y;\phi)+\sum_{z\in D_P}\ell_P(z;\phi).\]

Reward remains reward; preference remains preference; their noise models can differ. The hidden commitment is that a sufficiently shared \(\phi\) exists. If the score measures task success while the preference reflects style or safety, this assumption may fail.

Optimization fusion

At policy level, both datasets must participate in one maintained update:

\[(D_R,D_P)\longrightarrow g(\theta)\longrightarrow\theta'.\]

This need not mean a single loss. It means neither channel lives in a disconnected policy pipeline.

Operational fusion

A common learning state may improve data reuse, uncertainty accounting, auditing, reweighting, rollback, and the reuse of old feedback after policy changes. More importantly, it should expose disagreement. If rewards say \(A>B\) while people repeatedly report \(B\succ A\), a fused system should represent the conflict rather than quietly average it away.

numerical measurement → shared learner ← relational measurement

In this note, I use fusion for the stronger goal: preserve modality-specific information where possible, but let heterogeneous feedback participate in one coherent learning and control process.

3. Bandits are the clean test bed

Bandits strip away much of the temporal credit-assignment problem, making measurement heterogeneity easier to see. Feedback attaches directly to an arm or pair; long-horizon temporal structure is mostly absent.

Fusing Reward and Dueling Feedback in Stochastic Bandits makes this concrete. In its setting, both reward and dueling feedback are gathered each round; the learner is not choosing which modality to purchase. Elimination Fusion explores arms through both channels and shares a candidate set: evidence from either can remove a candidate. Decomposition Fusion partitions arms according to which channel yields the more favorable arm-dependent difficulty, then randomly assigns one feedback type to exploration and the other to exploitation. Under its stated assumption, the latter matches the regret lower bound up to a constant.

One way to read this result is that fusion lets each arm benefit from the more informative channel without requiring reward means and pairwise probabilities to be numerically interchangeable. Its objective is bandit regret efficiency, not a generic heterogeneous-feedback training interface. “Can both channels improve regret?” and “Can heterogeneous feedback live natively inside one learning paradigm?” are related but different questions.

In a bandit, likelihoods, confidence bounds, elimination, or allocation can often carry the fusion. Reinforcement learning adds another problem: feedback must not only be interpreted; it must be assigned back to decisions across time.

4. Policy gradient changes the problem

\[\boxed{\text{measurement heterogeneity}+\text{credit-assignment heterogeneity}}\]

Reward-based policy gradient eventually turns scalar outcomes into returns or advantages:

\[\nabla_\theta J(\theta)=\mathbb E\!\left[\sum_t A_t\nabla_\theta\log\pi_\theta(a_t\mid s_t)\right].\]

A preference such as \(\tau^+\succ\tau^-\) supplies no \(A_t\) for any action. The learner must decide what latent quantity generated the judgment, which decisions caused it, and how the relational observation becomes a policy update.

Preference → reward → policy

Classical RLHF provides one answer: fit \(\widehat r\) from comparisons, then optimize that scalar interface with RL. InstructGPT is a prominent instance of this pipeline. It is powerful because existing policy-optimization machinery already understands rewards. But it inserts an inferred object between observation and policy:

preferences → reward model → policy optimization

The reward model can be misspecified, become inaccurate under policy-induced distribution shift, or be exploited where its errors are largest. These are reasons to monitor the interface, not reasons to deny its usefulness.

Preference → policy

DPO is not itself a policy-gradient method; I include it here because it exposes a different interface by which preference data can reach policy parameters. Under its parameterization of the standard KL-regularized preference problem, \(D_P\rightarrow\pi\) can be implemented as a classification loss without an explicit reward-model training stage. This is an elegant simplification, but it does not by itself solve arbitrary long-horizon credit assignment.

For some LLM alignment formulations, a whole response behaves approximately like one structured action followed by sequence-level feedback—relatively bandit-like. An interactive agent instead produces \(s_0,a_0,s_1,a_1,\ldots,s_T\). Preferences may concern the whole trajectory while environment rewards arrive locally. The desired interface \((D_R,D_P)\rightarrow\pi\) must then solve both measurement and temporal credit.

5. What should a fused policy learner preserve?

The following are proposed desiderata, not accepted theory.

  1. Measurement preservation. Numerical rewards should not automatically be reduced to ranks, and preferences should not pretend to be absolute scores. If \(10>9\) and \(100>-100\) are compressed equally, the method should justify why that magnitude is irrelevant.
  2. A single policy-update interface. Both \(D_R\) and \(D_P\) should coherently affect the maintained \(g(\theta)\), even if they use different losses or estimators.
  3. Modality-specific observation models. \(P_R(y\mid\cdot)\) and \(P_P(z\mid\cdot)\) may have different noise, calibration, and evaluator effects. Sharing a representation does not make the channels statistically identical.
  4. Credit assignment. A trajectory preference must influence the decisions that generated it. Sequence-level order is not a per-action advantage.
  5. Conflict robustness. When channels disagree, a method might detect incompatibility, retain uncertainty, change weights, use modality-specific latent components, or decline to share some parameters. Blind averaging is not a resolution.
  6. Single-channel consistency. With \(D_P=\varnothing\), the method should reduce to sensible reward learning; with \(D_R=\varnothing\), to sensible preference learning.

These criteria let me examine a concrete method without asking only whether its benchmark score improved.

6. A concrete attempt: Dual-Feedback Actor

Fusing Rewards and Preferences in Reinforcement Learning proposes Dual-Feedback Actor (DFA). It is genuinely relevant because rewards and preferences enter one off-policy training process without first fitting a separate reward model. Numerical rewards train Q-networks from replay. Human preferences can be state-wise or trajectory-wise. When rewards are available, the method samples actions and synthesizes preference pairs using their learned Q-values; the actor is updated by a preference likelihood based directly on policy log-probabilities.

environment rewards → Q / critic → preference pairs → actor ← human preferences

In the paper’s state-wise tabular analysis, when preferences follow a Bradley–Terry model on the soft-optimal \(Q^\star\), the preference-loss minimizer is the Gibbs policy and coincides with the entropy-regularized SAC optimum when \(\lambda=\alpha/\beta\). This supplies an explicit bridge between preference optimization and an exploration-capable stochastic policy, and replay allows old transitions to be reused.

What is preserved—and what is compressed?

Numerical reward is not discarded: it shapes the critic. Yet the actor-facing fusion interface can turn critic values into an ordering, schematically

\[Q(s,a)>Q(s,a')\quad\Longrightarrow\quad a\succ a'.\]

Thus \(Q_A=10,Q_B=9\) and \(Q_A=100,Q_B=-100\) can induce the same synthetic label. From the measurement perspective above, part of the magnitude is compressed on its route to the actor. Under the theorem’s assumptions this construction has a clean justification; outside them, it is worth asking whether this is useful regularization or lost actor-facing information.

The compatibility bridge

The theoretical elegance comes from relating preference odds to the same entropy-regularized values that determine the optimal policy. In that model, rewards and preferences are views of a shared latent object. If human judgment depends on safety, style, delayed consequences, heterogeneous criteria, or presentation effects absent from reward Q-values, one might instead need \(Q_R(s,a)\neq Q_P(s,a)\). The theorem does not claim to cover that case; this is the boundary I am interested in.

Measurement model and decision model

DFA models preference probability through policy probabilities, using a temperature-like exponent \(\alpha\). The paper provides a theoretical reason for this coupling under its assumptions. Still, the construction raises a broader question: outside that model, should an evaluator’s observation process and a learner’s decision rule share exactly the same parameterization? A more flexible fused system may need to distinguish measurement model from decision model.

Credit and disagreement

DFA handles state-wise preferences directly and extends its policy-log-probability construction to trajectory comparisons. That makes it broader than a purely contextual response optimizer. But for long-horizon agents, how trajectory evidence should receive local credit remains a substantive modeling choice. Likewise, its theoretical alignment result relies on substantial compatibility between reward-generated and human preferences rather than testing that compatibility.

My assessment is therefore balanced: DFA demonstrates that reward and preference data need not live in separate pipelines. Its unification works partly by mapping both into a common actor-facing preference structure and assuming a strong relationship between their latent values. That makes it an excellent example of fusion—and a useful place to ask what more measurement-preserving fusion would require.

7. Toward native-feedback policy optimization

A. Shared latent-value fusion

One direction is to retain each observation model while letting both constrain a shared \(Q_\phi\):

\[D_R\to\mathcal L_R(Q_\phi),\quad D_P\to\mathcal L_P(Q_\phi),\qquad\mathcal L_{\mathrm{fuse}}=\lambda_R\mathcal L_R+\lambda_P\mathcal L_P.\]

If both terms are negative log-likelihoods from one shared generative model, the natural operation is first to sum observations. Introducing \(\lambda_R\) and \(\lambda_P\) then tempers or reweights evidence; those weights need a meaning—perhaps reliability, acquisition cost, or acknowledged misspecification—rather than serving as free tuning knobs. Rewards retain magnitude, preferences retain comparative likelihood, and policy improvement uses the shared representation. The immediate danger is forced sharing. A serious version needs compatibility tests, uncertainty, partial sharing, or hierarchical modality-specific components when one \(Q_\phi\) is indefensible.

B. Gradient-level fusion

Another direction constructs \(g_R(\theta)\) from reward data and \(g_P(\theta)\) from preference data:

\[g_{\mathrm{fuse}}=w_Rg_R+w_Pg_P.\]

Weights might reflect estimator variance, feedback reliability, acquisition cost, disagreement, off-policy mismatch, or credit quality. Yet even if \(g_R\simeq\nabla J_R\) and \(g_P\simeq\nabla J_P\), their weighted sum is not automatically an unbiased policy gradient for a fixed objective. If the weights change with \(\theta\), uncertainty, or disagreement, “what is being optimized?” becomes part of the theory, not an implementation detail.

The question I want to carry forward is concrete: can we construct policy-gradient estimators that let rewards remain rewards and preferences remain preferences, while allowing every observation to update one policy in a statistically coherent and conflict-aware way?

Conclusion

The opening score and comparison were not two formats for the same label. They were different measurements. Converting one into the other is a useful kind of fusion, but not the only kind. Bandits show that heterogeneous evidence can sometimes share a statistical object cleanly. Policy optimization makes the problem harder because credit assignment, replay, and disagreement enter at the same time.

One learner does not require one kind of feedback.

The challenge is to make heterogeneity usable without making it disappear.

Postscript: a later piece of evidence

The note above was written in March 2026. Best Arm Identification in Generalized Linear Bandits via Hybrid Feedback appeared in May 2026, so it was not evidence used to form the argument. I include it only because it later formalized a particularly clean version of statistical fusion: reward and arm-pair observations contribute to a shared likelihood-based confidence framework, while a joint design allocates measurements to certify the best arm.

References

← Back to Artifacts & Blogs
研究笔记

两种测量,
一个 Learner

Reward 把一个选项放到数值尺度上衡量;preference 把选项彼此比较。当 learner 同时收到二者,它如何使用全部信息,却不让任何一种信息消失?

郭道远 2026 年 3 月 约 15 分钟

一位 evaluator 给回答 A 打 8 分(满分 10 分);另一位同时看到 A 与 B 后选择了 B。这两条 observation 究竟告诉了我们什么?

二者都在判断“好”,却不可互换。分数把 A 放在一个人的数值尺度上;比较把 B 放在 A 的相对位置上。它没有告诉我们相差多少,而分数也没有说明第二个人会怎样使用同一量表。

这正是我认为 “同时使用 rewards 与 preferences” 比表面上困难的原因。如果系统已经收到两类观测,fusion 应该意味着什么?我的工作论点是:fusion 不应只是把所有 signal 转换成同一种 label。更严格的目标,是保留每种 measurement 所携带的信息,同时让全部证据进入同一个连贯的学习与控制过程。

1. 对“好”的两种测量

从 stochastic bandit 开始。设 arm \(i\) 的潜在价值为 \(u_i\)。数值 reward 是与单个 arm 关联的 outcome:

\[Y_i\sim P_R(\,\cdot\mid u_i),\qquad\text{例如 }Y_i=u_i+\varepsilon_i.\]

Pairwise label 则是与两个 arms 关联的 relational outcome:

\[Z_{ij}\sim\operatorname{Bernoulli}\!\left(\mu_P(u_i-u_j)\right).\]

单个 \(Z_{ij}\) 并不直接给出数值效用差;只有指定链接函数 \(\mu_P\) 并积累足够多的比较后,差值才可能被推断。Reward learner 能看到观测模型保留下来的数值幅度;preference 通道直接看到的则是一次带噪的排序事件。

这里还有一层容易被忽略。原始关系 \(i\succ j\) 只约束顺序;写成 \(\Pr(i\succ j)=\sigma(\beta(u_i-u_j))\),则额外规定了潜在差值如何变成选择概率。如果 \(\beta\) 未知,效用尺度与 \(\beta\) 可以相互抵消。定量结构来自参数模型,而不是来自 preference 这种观测格式本身。

不变性进一步说明差别。在常见 difference model 中,把所有效用变为 \(u_i+c\) 不会改变 preference probability:

\[(u_i+c)-(u_j+c)=u_i-u_j.\]

Reward observation 则可能保留 comparison 消去的 level information。Reward 把选项锚定到尺度;preference 把一个选项锚定到另一个。

令 \(u_i=x_i^\top\phi\),差异便成为几何问题:

\[Y_i\sim P_R(\cdot\mid x_i^\top\phi),\qquad Z_{ij}\sim P_P\!\left(\cdot\mid(x_i-x_j)^\top\phi\right).\]

Reward 沿 \(x_i\) 测量,preference 沿 \(x_i-x_j\) 测量。它们并非同一观测的两份带噪副本,而是作用于某个潜在对象的不同观测算子;这个对象可能共享,也可能并不共享。

2. Fusion 应该意味着什么?

我把问题整理成四个逐渐更强的层次;这只是本文的分类,不是公认术语。

Format fusion

把异质观测转为同一表示,例如 \(r_A>r_B\Rightarrow A\succ B\)。它很方便,有时也合理;但 \(10>9\) 与 \(100>-100\) 都变成同一个二元顺序,数值幅度可能被压缩。

Statistical fusion

保留不同观测模型,同时让二者约束同一个潜在对象:

\[\ell(\phi)=\sum_{y\in D_R}\ell_R(y;\phi)+\sum_{z\in D_P}\ell_P(z;\phi).\]

Reward 仍是 reward,preference 仍是 preference,噪声模型可以不同。隐藏前提是足够共享的 \(\phi\) 确实存在;若分数衡量任务完成度,而 preference 反映风格或安全性,这一假设就可能失败。

Optimization fusion

在 policy 层,两份数据必须共同参与维护中的 update:

\[(D_R,D_P)\longrightarrow g(\theta)\longrightarrow\theta'.\]

这不要求只有一个 loss,而是不能让两种 feedback 停留在互不相连的 policy pipeline。

Operational fusion

共同维护的学习状态可能改善数据复用、不确定性核算、审计、重加权与回滚,也便于在 policy 更新后继续使用旧反馈。更重要的是,它应把分歧暴露出来:若 rewards 指向 \(A>B\),人却持续报告 \(B\succ A\),系统不应悄悄把冲突平均掉。

numerical measurement → shared learner ← relational measurement

本文所说的 fusion 是一个更强的目标:尽可能保留各通道独有的信息,同时让异质 feedback 进入同一个连贯的学习与控制过程。

3. Bandit 是干净的试验场

Bandit 去掉了大部分跨时间的 credit assignment,使 measurement heterogeneity 更容易看清:feedback 直接对应某个 arm 或 pair,几乎没有长时域结构横在中间。

Fusing Reward and Dueling Feedback in Stochastic Bandits 中,每轮会同时收集 reward 与 dueling feedback;learner 并不主动二选一。Elimination Fusion 用两通道探索并共享 candidate set,任一通道都可淘汰候选。Decomposition Fusion 按 arm-dependent difficulty 把 arms 分给更有利的通道,再随机指定一种 feedback 用于 exploration、另一种用于 exploitation。在论文假设下,后者在常数范围内匹配 regret lower bound。

一种理解是:每个 arm 都能使用对自己更有信息量的通道,而无需把 reward mean 与 pairwise probability 当作可直接互换的数值。论文解决的是 bandit regret 的效率,并非通用的 heterogeneous-feedback training interface。“两通道能否改善 regret”与“所有异质 feedback 能否原生进入同一种学习范式”相关,却不是同一个问题。

在 bandit 中,likelihood、confidence、elimination 或 allocation 往往足以承载 fusion。RL 又增加一层:feedback 不只要被解释,还要跨时间分配回产生它的 decisions。

4. Policy gradient 改变了问题

\[\boxed{\text{measurement heterogeneity}+\text{credit-assignment heterogeneity}}\]

Reward-based policy gradient 最终把 scalar outcome 变成 return 或 advantage:

\[\nabla_\theta J(\theta)=\mathbb E\!\left[\sum_t A_t\nabla_\theta\log\pi_\theta(a_t\mid s_t)\right].\]

而 \(\tau^+\succ\tau^-\) 没有为任何 action 直接给出 \(A_t\)。Learner 必须回答:什么 latent quantity 产生了 judgment?哪些 decisions 导致它?relational observation 如何成为 policy update?

Preference → reward → policy

经典 RLHF 先从比较数据拟合 \(\widehat r\),再用 RL 优化这个标量接口;InstructGPT 是代表性实例。它的优势很实际:已有的 policy optimization 已经知道怎样处理 reward。但在观测与 policy 之间,也因此多了一个推断出来的对象:

preferences → reward model → policy optimization

Reward model 可能设定错误,在 policy 引起的分布变化下失准,也可能被 policy 钻到误差最大的区域。这些是需要监控这一接口的理由,并不否定它的价值。

Preference → policy

DPO 本身不是 policy-gradient 方法;我在这里讨论它,是因为它展示了 preference data 抵达 policy 参数的另一种接口。在其对标准 KL-regularized preference problem 的参数化下,\(D_P\rightarrow\pi\) 可化为 classification loss,无需显式训练 reward model。但它本身不解决任意长时域任务中的 credit assignment。

某些 LLM alignment formulation 把完整 response 近似视为一个 structured action 加 sequence-level feedback,比较接近 bandit。Interactive agent 则产生 \(s_0,a_0,s_1,a_1,\ldots,s_T\):preference 可能评价整条 trajectory,environment rewards 却在不同时间局部出现。此时 \((D_R,D_P)\rightarrow\pi\) 必须同时解决 measurement 与 temporal credit。

5. Fused policy learner 应保留什么?

以下六项是我提出的 desiderata,并非现成理论。

  1. Measurement preservation:数值 reward 不应默认退化为排序,preference 也不应伪装成绝对分数。若 \(10>9\) 与 \(100>-100\) 被同等压缩,方法需要说明数值幅度为何无关。
  2. Single policy-update interface:\(D_R\) 与 \(D_P\) 都应连贯影响维护中的 \(g(\theta)\),即使它们使用不同 loss 或 estimator。
  3. Modality-specific models:\(P_R(y\mid\cdot)\) 与 \(P_P(z\mid\cdot)\) 可以有不同的噪声、校准方式和评估者效应。共享表示不等于统计同质。
  4. Credit assignment:trajectory preference 必须影响产生它的 decisions;sequence-level order 不是 per-action advantage。
  5. Conflict robustness:通道冲突时,方法可以检测不相容、保留不确定性、调整权重、使用各通道自己的潜在分量,或者拒绝共享部分参数。简单取平均并没有解决冲突。
  6. Single-channel consistency:当 \(D_P=\varnothing\) 时应退化为合理 reward learning;当 \(D_R=\varnothing\) 时应退化为合理 preference learning。

这组标准让我们能分析具体方法,而不只是问 benchmark score 是否提高。

6. 一次具体尝试:Dual-Feedback Actor

Fusing Rewards and Preferences in Reinforcement Learning 提出 Dual-Feedback Actor(DFA)。它直接相关,因为 rewards 与 preferences 进入同一个 off-policy training process,不需先拟合独立 reward model。Numerical rewards 从 replay 训练 Q-networks;human preferences 可为 state-wise 或 trajectory-wise。当 reward 可用时,方法采样 actions,用学得的 Q-values 合成 preference pairs;actor 则由基于 policy log-probability 的 preference likelihood 更新。

environment rewards → Q / critic → preference pairs → actor ← human preferences

更准确地说,在论文的 state-wise tabular analysis 中,如果 preference 遵循定义在 soft-optimal \(Q^\star\) 上的 Bradley–Terry model,那么 preference loss 的唯一 minimizer 是 Gibbs policy;当 \(\lambda=\alpha/\beta\) 时,它与 entropy-regularized SAC optimum 一致。这在 preference optimization 与具有随机探索能力的 policy 之间建立了明确桥梁,replay 也使旧 transitions 可以复用。

什么被保留,什么被压缩?

数值 reward 没有被丢弃:它塑造 critic。但面向 actor 的 fusion 接口会把 critic values 转成排序,例如

\[Q(s,a)>Q(s,a')\quad\Longrightarrow\quad a\succ a'.\]

因此 \(Q_A=10,Q_B=9\) 与 \(Q_A=100,Q_B=-100\) 可能生成相同的 synthetic label。从 measurement 角度看,部分数值幅度在到达 actor 的路上被压缩。在 theorem assumptions 内,这样做有清楚依据;离开这些假设后,则值得追问它是有效的 regularization,还是 actor 端的信息损失。

Compatibility bridge

理论优雅之处,是把 preference odds 与决定 optimal policy 的同一组 entropy-regularized values 联系起来;在此模型中,rewards 与 preferences 是同一个潜在对象的两种观测。若人的判断还依赖 reward Q-values 未表示的安全性、风格、延迟后果、异质标准或展示效应,则可能需要 \(Q_R(s,a)\neq Q_P(s,a)\)。论文 theorem 没有声称覆盖此情形;这正是我关心的研究边界。

Measurement model 与 decision model

DFA 用 policy probability 及指数 \(\alpha\) 建模 preference probability。在论文假设下,这种 coupling 有理论理由。但更一般地,评估者的观测过程与 learner 的决策规则是否应共享完全相同的参数化?更灵活的 fused system 可能需要把 measurement model 与 decision model 分开。

Credit 与 disagreement

DFA 直接处理 state-wise preference,并把基于 policy log-probability 的构造扩展到 trajectory comparisons,因此比纯 contextual response optimizer 更广。但对 long-horizon agent,如何把整条 trajectory 的评价局部地归因给沿途决策,仍是重要选择。更准确地说,依赖强相容性的是它的理论对齐结论,而算法本身并不会先检验 reward-generated 与 human preferences 是否相容。

因此我的判断是平衡的:DFA 证明了 reward 与 preference data 不必分居两套 pipeline;但其统一部分依靠把二者映射到共同 actor-facing preference structure,并假定 latent values 有强关系。它既是很好的 fusion 实例,也是追问 measurement-preserving fusion 的合适起点。

7. 走向 native-feedback policy optimization

A. Shared latent-value fusion

一种方向是保留各自的观测模型,同时约束共享的 \(Q_\phi\):

\[D_R\to\mathcal L_R(Q_\phi),\quad D_P\to\mathcal L_P(Q_\phi),\qquad\mathcal L_{\mathrm{fuse}}=\lambda_R\mathcal L_R+\lambda_P\mathcal L_P.\]

如果两项确实来自同一个生成模型,那么最自然的统计操作首先是把各条观测的 log-likelihood 相加。额外加入 \(\lambda_R\) 和 \(\lambda_P\),实际上是在 temper 或重加权证据;这些权重应当对应可靠性、采集成本或已知的模型偏差,而不只是可调超参数。Reward 保留数值幅度,preference 保留比较 likelihood,policy improvement 使用共享表示。危险在于强行共享:严肃版本需要相容性检验、不确定性、部分共享,或各通道自己的层级分量。

B. Gradient-level fusion

另一方向分别构造 reward gradient \(g_R(\theta)\) 与 preference gradient \(g_P(\theta)\):

\[g_{\mathrm{fuse}}=w_Rg_R+w_Pg_P.\]

权重可以考虑 estimator variance、反馈可靠性、采集成本、通道分歧、off-policy mismatch 或 credit assignment 的质量。但即使 \(g_R\simeq\nabla J_R\)、\(g_P\simeq\nabla J_P\),加权和也不自动成为某个固定目标的无偏 policy gradient。尤其当权重随 \(\theta\)、不确定性或分歧变化时,“究竟在优化什么”本身就是理论问题,而不是实现细节。

我想带走的具体问题是:能否构造 policy-gradient estimators,让 rewards 仍是 rewards、preferences 仍是 preferences,同时让每条 observation 以 statistically coherent、conflict-aware 的方式更新同一个 policy?

结论

开头的分数与比较不是同一 label 的两种格式,而是两种不同的 measurement。把一种转成另一种是有用的 fusion,却不是唯一可能。Bandit 表明异质证据有时可以干净地共享一个统计对象;到了 policy optimization,credit assignment、replay 与通道分歧同时出现,问题便难得多。

One learner 不要求只有一种 feedback。

挑战是利用这种差异,而不是抹掉它。

后记:后来出现的一项证据

以上笔记写于 2026 年 3 月。 Best Arm Identification in Generalized Linear Bandits via Hybrid Feedback 于 2026 年 5 月出现,因此不是形成原论证时使用的证据。这里只记录它后来形式化了一个很干净的 statistical fusion:reward 与 arm-pair observations 进入共享的 likelihood-based confidence framework,再通过 joint design 分配测量以认证 best arm。

参考资料

← 返回 Artifacts & Blogs