GLM 5/5.1 · Chip comparison

GLM 5/5.1 — GB300 NVL72 vs MI355X

Head-to-head AI inference benchmark comparison of GB300 NVL72 (NVIDIA Blackwell) and MI355X (AMD CDNA 4) on GLM 5/5.1. Latency, throughput, and cost across LLM workloads. Use the chart controls below to switch sequences, precisions, and metrics — same interactions as the main inference chart.

Setting 30 tok/s/user as the target on GLM 5/5.1, GB300 NVL72 produces 11341 tok/s/chip ($0.06 per million tokens) and MI355X produces 907 ($0.46). GB300 NVL72 is 715% cheaper per token; GB300 NVL72 delivers 1150% more tok/s/chip.

At 42 tok/s/user interactivity on GLM 5/5.1, GB300 NVL72 delivers 10234 tok/s/chip at $0.06 per million tokens; MI355X delivers 452 tok/s/chip at $0.91. GB300 NVL72 is 1350% cheaper per token; GB300 NVL72 delivers 2166% more tok/s/chip at this point.

GB300 NVL72 posts 8633 tok/s/chip for $0.07 per million tokens at 55 tok/s/user on GLM 5/5.1; MI355X posts 228 tok/s/chip for $1.83. GB300 NVL72 is 2364% cheaper per token; GB300 NVL72 delivers 3686% more tok/s/chip. (Numbers reflect the default 1k/1k · fp4 selection for this URL — table and chart below update if you change sequence, precision, or model in the controls.)

View performance-per-dollar view →

Interpolated from real benchmark data. Edit target interactivity values below to compare at different operating points.
Metric
Interactivity (tok/s/user)
Interactivity (tok/s/user)
Interactivity (tok/s/user)
Throughput (tok/s/chip)
GB300 NVL72:11341.4MI355X:907.1
GB300 NVL72:10233.9MI355X:451.6
GB300 NVL72:8632.9MI355X:228.0
Cost ($/M tok)
GB300 NVL72:$0.056MI355X:$0.460
GB300 NVL72:$0.063MI355X:$0.909
GB300 NVL72:$0.074MI355X:$1.833
tok/s/MW
GB300 NVL72:5349707MI355X:434038
GB300 NVL72:4827317MI355X:216095
GB300 NVL72:4072116MI355X:109104
Concurrency
GB300 NVL72:~4301MI355X:~62
GB300 NVL72:~3770MI355X:~22
GB300 NVL72:~1946MI355X:~9

Inference Performance

Inference performance metrics across different models, hardware configurations, and serving parameters.

Vendor:
Deployment:
Spec Decoding: