GB300 NVL72 FP4: MTP vs Off Speculative Decoding
Speculative decoding comparison of MTP versus Off on GB300 NVL72 FP4 (NVIDIA Blackwell) running Qwen 3.5 397B-A17B. Throughput, cost, and interactivity differences across LLM workloads. Use the chart controls below to switch sequences and metrics — same interactions as the main inference chart.
MTP acceptance-rate implementations differ across inference engines. Points from different engines are not directly comparable on the same curve — throughput and cost at matched interactivity may reflect engine-level differences rather than pure speculative decoding gains. Interpret cross-engine comparisons with caution.
Off hits 6025 tok/s/GPU for $0.12 per million tokens at 90 tok/s/user on Qwen 3.5 397B-A17B (GB300 NVL72 FP4). No MTP data at this operating point.
Off: 2801 tok/s/GPU, $0.26 per million tokens at 136 tok/s/user on Qwen 3.5 397B-A17B (GB300 NVL72 FP4). MTP is unmeasured here.
At 183 tok/s/user on Qwen 3.5 397B-A17B (GB300 NVL72 FP4), Off delivers 1338 tok/s/GPU at $0.55 per million tokens; MTP hasn't been benchmarked at this target. (Numbers reflect this URL's pinned 8k/1k · fp4 workload — changing sequence or model updates both the table and chart; the table stays pinned to this page's precision, so precision toggles in the controls affect the chart only.)

| Metric | Interactivity (tok/s/user) | Interactivity (tok/s/user) | Interactivity (tok/s/user) |
|---|---|---|---|
| Throughput (tok/s/gpu) | MTP:—Off:6024.7 | MTP:—Off:2801.0 | MTP:—Off:1338.4 |
| Cost ($/M tok) | MTP:—Off:$0.122 | MTP:—Off:$0.256 | MTP:—Off:$0.553 |
| tok/s/MW | MTP:—Off:2841855 | MTP:—Off:1321248 | MTP:—Off:631330 |
| Concurrency | MTP:—Off:~71 | MTP:—Off:~21 | MTP:—Off:~7 |
Inference Performance
Inference performance metrics across different models, hardware configurations, and serving parameters.