The problem
A 35B mixture-of-experts model at Q4 is about 20 GB of weights, and the card in this machine holds 16. Buying more card is not the interesting answer for an operator whose inference has to sit next to the data. The interesting answer is that routed experts are touched sparsely: a token uses 8 of 256 per layer, so the working set of a decode step is a fraction of the model. Everything hinges on what happens when the expert you need is not in VRAM.
What we built
titan-engine serves the model with a tiered expert scheme: a fraction of every layer's experts stays resident in VRAM, the rest live in host memory, and the ones that miss are computed on the CPU in the same step, with no extra GPU round trip and no host synchronisation added. The measured budget for a 35B decode step at MTP=2 is 24.0 ms for 2.70 tokens, of which 14.5 ms is GPU kernel time and about 4.4 ms is the CPU expert pass plus its syncs.
- Every kernel in the path is ours: written as Rust that emits PTX and hosted on NVIDIA's cuda-oxide, so the build needs no nvcc. GGUF dequantisation, Q8_0 dense projections, the expert gemvs and the prefill MMQ tiles all come from that tree.
- The CPU expert pass has its own fork-join pool with workers that spin rather than park, because a decode step issues two short passes per MoE layer with GPU work in between and waking rayon's pool costs more than the pass.
- Bit-identity is a design rule where it can be: the CPU twin emulates the GPU kernel's 128-lane reduction so a tiered forward and an all-resident forward produce the same text, and every kernel port is gated against llama.cpp's outputs.
- Multi-token decode (MTP=2), a hybrid prefix cache, chunked prefill and CUDA graphs for the batch-1 decode segments are all in the deployed binary.
Where the time actually goes
The standing benchmark suite runs three tiers: the matmul kernels against llama.cpp on identical device buffers, a path tier that takes one MoE block's forward out of a real profile, and end-to-end decode, prefill and time-to-first-token. It has paid for itself by killing three of our own ideas:
- AVX-VNNI for the CPU expert dots is implemented, bit-identical and merged — and worth nothing end to end. It is 1.2–1.26× on Q6_K rows and 1.02–1.06× on Q4_K and Q5_K in isolation; measured through the service it moved 80B decode from 14.2 to 14.1 tok/s. The CPU pass is dispatch-bound, not dot-bound: the kernel streams 10–17 GB/s on a box whose measured read ceiling is 47 GB/s.
- A learned expert-eviction policy was simulated against the static, LRU and LFU placements and landed within ±1.6% of the best of them, against a bar of +3%. Not shipped.
- CUDA graphs gave +2.6–2.9% on decode, under the bar, and are off in the service while one gate is still open.
What it buys, honestly
Two things are real and measured. Prefill expert streaming: a layer whose chunk is large enough streams every expert it needs through a ring of device staging buffers on a copy stream, and a 13k-token prompt went from 336 to 1350 tok/s. And the model-level picture: a 24.0 ms step against 19.6 ms with the CPU pass disabled, which is 137.9 tok/s the engine cannot reach on this hardware without holding the whole model in VRAM.
The honest other half: the kernels that lag llama.cpp most are the routed-expert gemvs at decode (1.2–2.0×) and the MTP verify batch (1.5–2.6×), which together are about 28% of service GPU time. The fix is identified down to the kernel — llama.cpp's row-per-expert grid — and not yet in the tiered path.