These are notes for this paper.

Context

MoE was first introduced in 2017 and within the last few years there has been steady rise in deploying this algorithm.

This work is expands on previous works

  • Sparse Upcycling (Komatsuzaki et al., 2023)
  • Brain-Train-Mix (BTX) (Sukhbaatar et al., 2024)
  • Branch-Train-Merge (BTM) (Li et al., 2022)

Novel approach utilized in Nexus is using a dynamic router within MoE

  • Usual MoE routers have fixed number of experts
  • Usual MoE only route using the tokens
  • Nexus router uses domain embeddings and expert embedding for routing hence new MoEs could be added.

Sparse Upcycling

  • Training mixture-of-experts from dense checkpoints
  • Take a dense model and make it MoE by only training the router
  • Keep the transformer bit intact, copy the FFN part, add multiples of it with a router
Sparse UpCycling

Branch-Train-Mix (BTX)

  • Training mixture-of-experts from dense checkpoints
  • Take a dense model and make it MoE by only training the router
  • Keep the transformer bit intact, copy the FFN part, add multiples of it with a router
  • Brain-Train-Merge: Averages everything and no routers.
BTX

Nexus

Domain embeddings are utilized with SwiGLU to generate the expert embeddings which later on determine which router should be used given the tokens. This approach is a kin to:

  1. Mahabadi et al., 2021 (Parameter-efficient multi-task fine-tuning)
  2. Üstün et al., 2022 (Hyper-X)

$e_i = P_r(d_i)$ (Domain to Expert Embeddings)

$= W_2 \cdot \text{SwiGLU}(W_1 \cdot d_i)$

Note: Similar to DeepSeek V3, it keeps a shared FFN.

Nexus
  • Mixing expert lms into a mixture-of-experts lm
  • Merges already expert LLMs
  • Takes their FFNs and averages transformer bits to create a shared body

Summary Table

Feature Sparse Upcycling BTM (Merge) BTX (Mix) & Nexus
Starting Point One general model Multiple specialized models Multiple specialized models
What it does with FFNs Copies them identically Averages them together Collects them as experts
Is there a Router? Yes, trained from scratch No router Yes, trained from scratch
Final Model Type Sparse MoE Dense Sparse MoE

Expermintal setup

  • Two model sizes: 470M, 2.8B
  • Five different categories
  • Trained an expert then added it to study properties
Results Across Domain

Results for Upcycling

Results Across Domain

Result: Expert selection

Results on Expert Selection

Ablations

  1. Effects of load balance weights
Load Balance Ablation
  1. Altering data training composition
  2. Effectiveness of domain embedding

embed_effectiveness</img>

Embedding Effectiveness