I Built a Router.
Did I Build an MoE Problem?
A Deep Learning course project taught me how to model Bayesian routing. A later Machine Learning Systems course changed the question: when does sparse routing become a systems property rather than an architectural detail?
This reflection grew out of my course project, Bayesian Applications in Mixture of Experts. The project was a useful routing testbed. It was not designed as a large-scale MoE serving system, and the distinction is the point of this article.
I thought the project was asking whether Bayesian routing could improve an MoE. Looking back, I had built a model of routing decisions—but not yet a model of their system consequences.
1. What the course project actually modeled
The model starts with a ResNet-18 backbone that maps an image to a 32-dimensional representation:
Sixteen small experts then map that representation to class logits. For CIFAR-100, each expert is a two-layer MLP:
The router selects the top two experts. The implementation dispatches only to experts appearing in that top-k set, so the sparsity is real rather than dense computation followed by masking. The Bayesian part belongs to the router: its weights and biases have factorized Gaussian posteriors. The expected variant uses posterior moments; the Monte Carlo variant samples router parameters and optimizes an ELBO-style objective.
As a Deep Learning course project, this was a sensible abstraction. It let us compare routing, calibration, expert balance, specialization, and OOD confidence without needing a cluster-scale implementation.
But the parameter structure reveals what the model cannot study. One CIFAR-100 expert has about 8,612 parameters; all sixteen contain about 138K. The shared ResNet-18 backbone, after replacing its final layer, has roughly 11.19M parameters. The entire expert pool is only about 1.2% of the model.
Only two experts are active per image, but sparsifying a small fraction of the model cannot reveal much about whole-system efficiency. Parameter share is not runtime share—the actual fraction must be measured with a profiler—but the architecture already tells us that this is primarily a routing-learning testbed, not an MoE scaling testbed.
2. Sparse routing existed, but sparsity barely mattered
The appeal of large MoE models comes from separating total capacity from active computation. Let Ps denote shared parameters, Pe one expert, E the number of experts, and k the number activated for one token:
With fixed k, increasing E increases total capacity without increasing active expert computation at the same rate. This regime motivates systems such as Switch Transformer.
Our model lived in almost the opposite regime. The expensive shared backbone ran for every image, while the sparsified experts formed a small classifier at the end. If expert computation occupies fraction ρ of dense inference time, an Amdahl-style idealized upper bound is:
When ρ is small, even a large E/k produces little end-to-end speedup. Sparse routing existed mathematically, but the project did not operate where sparsity was likely to matter as a systems property.
3. What scale changes
Taking Machine Learning Systems gave me a more useful way to read this architecture. Scaling the project would not merely make the same experiment larger. It would introduce different problems.
Routing becomes token-level and repeated
Our router makes one decision per image at one MoE layer. Transformer MoEs can route every token at multiple layers. This creates cross-layer expert paths, token-level imbalance, decode-time locality, and opportunities for prediction and prefetching that the course model cannot express.
Experts stop being local by default
The final comparison ran on one GPU. It therefore removed the dispatch and all-to-all communication that often dominate expert-parallel MoE systems. Once experts live across devices, routing also becomes a placement and communication decision.
The project modeled who should process an input. A large MoE system must also decide where that expert lives, whether its weights are resident, and what must move before computation begins.
4. Sparse compute is not efficient inference
MoE gives conditional computation. It does not automatically lower GPU memory use. If every expert remains resident, weight memory still scales with all experts:
Active computation contains k · Pe, but storage contains E · Pe. Sparse activation is not sparse storage. Turning it into memory savings requires sharding, quantization, caching, or offloading—the gap addressed by DeepSpeed-MoE and MoE-Infinity.
Latency has the same problem. Counting expert FLOPs sees only one term:
A small k reduces expert work, but wall-clock latency improves only if communication, transfer, and synchronization do not consume the saving. Algorithmic sparsity creates an opportunity. The system still has to realize it.
5. From a model objective to a systems objective
A systems formulation must choose a deployment objective and constraints. For example:
The router is no longer judged only by task loss. It participates in resource allocation under quality and memory constraints.
This changes how I would choose a research direction: profile first. If GPU memory is the bottleneck, study residency, quantization, caching, or offloading. If PCIe transfer dominates, study prediction and prefetch. If all-to-all communication dominates, study topology-aware routing, expert placement, replication, or overlap.
Starting with an algorithm and searching for a bottleneck it might solve is much easier than establishing that the bottleneck exists.
6. A possible next step: uncertainty as a resource signal
The Bayesian router can still matter, but perhaps not primarily because it changes accuracy or ECE. Its posterior can estimate the probability that an expert will appear in the top-k set:
If only B experts can be prefetched, choose the set with the largest posterior access probability:
Or choose the smallest resident set that covers the predicted top-k experts with probability at least 1 − δ:
Here δ controls a memory-versus-cache-miss trade-off. The same uncertainty could determine dynamic k: confident tokens use fewer experts; ambiguous tokens receive more compute.
This is a proposed direction. The repository does not implement expert caching, prefetching, or dynamic k, and its image-level single-layer router cannot validate the idea at realistic scale. But it gives the Bayesian component a systems role that grows naturally from the original project.
7. What changed in my question
The Deep Learning project taught me how to construct and compare routers. Machine Learning Systems taught me to ask whether an architecture creates a meaningful compute, memory, or communication opportunity—and whether an implementation captures the deployment bottleneck.
The course project studied routing quality. MoE at scale is a resource-allocation problem.
The question I would take forward is no longer simply, “Can Bayesian routing improve accuracy?” It is: can routing uncertainty help a system decide what to compute, what to keep in accelerator memory, and what to move before it is needed?
References and project material
- Course project repository
- Switch Transformers
- DeepSpeed-MoE
- MoE-Infinity
- Fate: Expert Prediction and Prefetching
我实现了一个 Router。
但它构成了真正的 MoE 问题吗?
Deep Learning 课程项目让我学会如何建模 Bayesian routing;后来 Machine Learning Systems 课程改变了我的问题:稀疏路由何时才从架构细节变成系统属性?
本文源于我的课程项目 Bayesian Applications in Mixture of Experts。它是一个有用的 routing testbed,但并不是面向大规模 MoE serving 的系统;这一区别正是本文的主题。
我原以为项目研究的是 Bayesian routing 能否改善 MoE。回头看,我建模了路由决策,却还没有建模路由决策带来的系统后果。
1. 课程项目实际建模了什么
模型先通过 ResNet-18 backbone 将图像映射到 32 维表示:
随后 16 个小型 expert 将表示映射到类别 logits。CIFAR-100 上每个 expert 是两层 MLP:
Router 选择 top-2 experts,而且实现确实只执行 top-k 集合中的 experts。Bayesian 部分只在 router:权重与 bias 使用 factorized Gaussian posterior;expected 版本使用 posterior moments,Monte Carlo 版本采样 router 参数并优化 ELBO-style objective。
作为 Deep Learning 课程项目,这是合理的抽象。但参数结构也说明了它不能研究什么:16 个 experts 合计约 0.138M 参数,而 shared backbone 约 11.19M;expert pool 只占模型约 1.2%。
参数占比不等于 runtime 占比,实际比例仍需 profiler 测量;但这个结构已说明它主要是 routing-learning testbed,而不是 MoE scaling testbed。
2. 存在稀疏路由,但稀疏性几乎不影响系统
大型 MoE 的吸引力来自 total capacity 与 active computation 的分离:
固定 \(k\) 并增加 \(E\),总容量增长,而每个 token 激活的 expert computation 不同比例增长。要让 sparsity 真正重要,expert pool 必须是模型的重要部分:
我们的模型几乎处在相反区间:昂贵的 shared backbone 对每张图都执行,稀疏 experts 只是末端的小分类器。若 expert computation 占 dense inference time 的比例为 \(\rho\),理想化的 Amdahl 上界为:
当 \(\rho\) 很小时,\(E/k\) 再大也难以带来明显的端到端加速。
3. Scale 会改变问题本身
Machine Learning Systems 课程让我用另一种方式理解这个架构。扩大项目规模并不是把相同实验做大,而是会引入新的问题。
路由会变成 token-level 且跨层重复
课程模型每张图只在一个 MoE layer 做一次决策;Transformer MoE 会在多个 layer 为每个 token 决策,从而产生 cross-layer paths、token-level imbalance、decode-time locality,以及 prediction 与 prefetch 机会。
Experts 不再默认位于本地
最终比较运行在单 GPU 上,因此没有 expert parallelism 中常见的 dispatch 与 all-to-all communication。Experts 分布到多个设备后,路由同时也是 placement 与 communication 决策。
项目决定谁处理输入;大型 MoE 系统还必须决定 expert 在哪里、权重是否 resident,以及计算前需要移动什么。
4. Sparse compute 不等于高效 inference
MoE 提供 conditional computation,但不会自动降低 GPU memory。如果所有 experts 都常驻设备,weight memory 仍随全部 experts 增长:
Active computation 只含 \(kP_e\),storage 却包含 \(EP_e\)。Sparse activation 不是 sparse storage;要获得显存收益,还需要 sharding、quantization、caching 或 offloading。
Latency 也一样。只计算 expert FLOPs 会遗漏多项成本:
小 \(k\) 只减少 expert work;只有 communication、transfer 与 synchronization 没有吞掉节省,wall-clock latency 才会改善。算法稀疏性只创造机会,系统必须把机会变成收益。
5. 从模型目标转向系统目标
系统研究必须先定义 deployment objective 与 constraints,例如:
Router 不再只由 task loss 判断,而成为质量与 memory constraints 下的资源分配策略。
这也改变了研究方向的选择顺序:先 profile。若 GPU memory 是瓶颈,研究 residency、quantization、caching 或 offloading;若 PCIe transfer 占主导,研究 prediction 与 prefetch;若 all-to-all communication 占主导,则研究 topology-aware routing、expert placement、replication 或 overlap。
6. 下一步可能方向:把 uncertainty 当作 resource signal
Bayesian router 的价值也许不只在 accuracy 或 ECE。Posterior 可以估计某个 expert 进入 top-k 的概率:
若只能额外 prefetch \(B\) 个 experts,可以选择 posterior access probability 最大的集合:
也可以寻找最小 resident set,使真正需要的 top-k experts 以至少 \(1-\delta\) 的概率全部在其中:
这里 \(\delta\) 直接控制 memory 与 cache-miss risk 的权衡。同一 uncertainty 也可用于 dynamic \(k\):确定的 token 使用更少 experts,模糊的 token 获得更多 compute。
这只是 proposed direction。仓库并未实现 expert caching、prefetching 或 dynamic \(k\),单层 image-level router 也不能在真实 MoE scale 上验证它。但这个方向为原项目中的 Bayesian component 找到了自然的 systems role。
7. 我的研究问题如何改变
Deep Learning 项目教会我构建与比较 routers;Machine Learning Systems 则让我先问:这个架构是否真的创造了 compute、memory 或 communication 机会?实现是否捕捉了 deployment 中真正的 bottleneck?
课程项目研究 routing quality;大规模 MoE 则是 resource-allocation problem。
我下一步想问的不再只是“Bayesian routing 能否提高 accuracy”,而是:routing uncertainty 能否帮助系统决定计算什么、哪些内容应留在 accelerator memory,以及哪些权重需要在使用前移动?