GB300 NVL72 FP8: MTP vs Off Speculative Decoding
Speculative decoding comparison of MTP versus Off on GB300 NVL72 FP8 (NVIDIA Blackwell) running Qwen 3.5 397B-A17B. Throughput, cost, and interactivity differences across LLM workloads. Use the chart controls below to switch sequences and metrics — same interactions as the main inference chart.
MTP acceptance-rate implementations differ across inference engines. Points from different engines are not directly comparable on the same curve — throughput and cost at matched interactivity may reflect engine-level differences rather than pure speculative decoding gains. Interpret cross-engine comparisons with caution.
MTP posts 10224 tok/s/GPU for $0.07 per million tokens at 93 tok/s/user on Qwen 3.5 397B-A17B (GB300 NVL72 FP8); Off posts 2800 tok/s/GPU for $0.26. MTP is 264% cheaper per token; MTP delivers 265% more tok/s/GPU. Draft-token acceptance rates determine whether speculative decoding helps or hurts at a given concurrency level.
Throughput at 131 tok/s/user on Qwen 3.5 397B-A17B (GB300 NVL72 FP8): MTP hits 5267 tok/s/GPU, Off hits 1529. Per-million costs land at $0.14 and $0.48 respectively. MTP is 236% cheaper per token; MTP delivers 244% more tok/s/GPU. Speculative decoding trades extra compute on draft tokens for fewer decoding steps — the payoff depends on sequence length and batch size.
Toward the upper edge of the 55–206 tok/s/user interactivity band, at 169 tok/s/user on Qwen 3.5 397B-A17B (GB300 NVL72 FP8): MTP runs 3181 tok/s/GPU at $0.23/M tokens, Off runs 728 at $1.01/M. MTP is 335% cheaper per token; MTP delivers 337% more tok/s/GPU. Gains from speculative decoding vary by workload; short-output prompts tend to benefit less. (Numbers reflect this URL's pinned 8k/1k · fp8 workload — changing sequence or model updates both the table and chart; the table stays pinned to this page's precision, so precision toggles in the controls affect the chart only.)

| Metric | Interactivity (tok/s/user) | Interactivity (tok/s/user) | Interactivity (tok/s/user) |
|---|---|---|---|
| Throughput (tok/s/gpu) | MTP:10223.9Off:2799.8 | MTP:5266.8Off:1529.2 | MTP:3180.7Off:727.8 |
| Cost ($/M tok) | MTP:$0.072Off:$0.262 | MTP:$0.143Off:$0.482 | MTP:$0.232Off:$1.008 |
| tok/s/MW | MTP:4822602Off:1320647 | MTP:2484317Off:721311 | MTP:1500330Off:343282 |
| Concurrency | MTP:~906Off:~30 | MTP:~122Off:~12 | MTP:~46Off:~4 |
Inference Performance
Inference performance metrics across different models, hardware configurations, and serving parameters.