DeepSeek R1 — B200 vs MI300X
Head-to-head AI inference benchmark comparison of B200 (NVIDIA Blackwell) and MI300X (AMD CDNA 3) on DeepSeek R1. Latency, throughput, and cost across LLM workloads. Use the chart controls below to switch sequences, precisions, and metrics — same interactions as the main inference chart.
Setting 36 tok/s/user as the target on DeepSeek R1, B200 produces 5320 tok/s/chip ($0.09 per million tokens) and MI300X produces 268 ($0.99). B200 is 976% cheaper per token; B200 delivers 1888% more tok/s/chip.
At 47 tok/s/user interactivity on DeepSeek R1, B200 delivers 4125 tok/s/chip at $0.12 per million tokens; MI300X delivers 186 tok/s/chip at $1.41. B200 is 1112% cheaper per token; B200 delivers 2113% more tok/s/chip at this point.
B200 posts 2816 tok/s/chip for $0.17 per million tokens at 58 tok/s/user on DeepSeek R1; MI300X posts 113 tok/s/chip for $2.32. B200 is 1279% cheaper per token; B200 delivers 2385% more tok/s/chip. (Numbers reflect the default 1k/1k · fp8 selection for this URL — table and chart below update if you change sequence, precision, or model in the controls.)
| Metric | Interactivity (tok/s/user) | Interactivity (tok/s/user) | Interactivity (tok/s/user) |
|---|---|---|---|
| Throughput (tok/s/chip) | B200:5320.1MI300X:267.6 | B200:4125.0MI300X:186.4 | B200:2816.0MI300X:113.3 |
| Cost ($/M tok) | B200:$0.092MI300X:$0.986 | B200:$0.116MI300X:$1.410 | B200:$0.168MI300X:$2.322 |
| tok/s/MW | B200:3111188MI300X:192553 | B200:2412259MI300X:134115 | B200:1646769MI300X:81533 |
| Concurrency | B200:~1947MI300X:~31 | B200:~1357MI300X:~17 | B200:~1050MI300X:~8 |
Inference Performance
Inference performance metrics across different models, hardware configurations, and serving parameters.