Intel SYCL
Bucket and reuse
At four or more tokens, the device groups selected pairs by expert and handles up to four tokens per weight pass. Smaller batches use a fused per-pair device route. Shape-aware workgroups use one subgroup for 512-row gate/up projections and four for the 2,048-row down projection.
NVIDIA CUDA
Stay on the quantized fast path
Selected-expert projections remain in device-routed MMVQ/MMQ kernels. On Blackwell with Q8_0 and twelve concurrent token columns, Treebeard uses the direct 12-column MMVQ path. Those twelve columns are tokens—not twelve experts.