gpt-oss 120B · Chip comparison

gpt-oss 120B — H100 vs MI325X

Head-to-head AI inference benchmark comparison of H100 (NVIDIA Hopper) and MI325X (AMD CDNA 3) on gpt-oss 120B. Latency, throughput, and cost across LLM workloads. Use the chart controls below to switch sequences, precisions, and metrics — same interactions as the main inference chart.

At 78 tok/s/user interactivity on gpt-oss 120B, H100 delivers 3838 tok/s/chip at $0.08 per million tokens; MI325X delivers 1385 tok/s/chip at $0.22. H100 is 168% cheaper per token; H100 delivers 177% more tok/s/chip at this point.

H100 posts 3493 tok/s/chip for $0.09 per million tokens at 90 tok/s/user on gpt-oss 120B; MI325X posts 1126 tok/s/chip for $0.27. H100 is 196% cheaper per token; H100 delivers 210% more tok/s/chip.

Throughput at 102 tok/s/user on gpt-oss 120B: H100 hits 3130 tok/s/chip, MI325X hits 771. Per-million costs land at $0.10 and $0.40 respectively. H100 is 281% cheaper per token; H100 delivers 306% more tok/s/chip. (Numbers reflect the default 1k/1k · fp4 selection for this URL — table and chart below update if you change sequence, precision, or model in the controls.)

View performance-per-dollar view →

Interpolated from real benchmark data. Edit target interactivity values below to compare at different operating points.
Metric
Interactivity (tok/s/user)
Interactivity (tok/s/user)
Interactivity (tok/s/user)
Throughput (tok/s/chip)
H100:3838.4MI325X:1385.3
H100:3492.5MI325X:1126.4
H100:3129.9MI325X:770.7
Cost ($/M tok)
H100:$0.084MI325X:$0.225
H100:$0.092MI325X:$0.272
H100:$0.104MI325X:$0.396
tok/s/MW
H100:2801737MI325X:819683
H100:2549295MI325X:666511
H100:2284627MI325X:456008
Concurrency
H100:~64MI325X:~23
H100:~64MI325X:~28
H100:~64MI325X:~16

Inference Performance

Inference performance metrics across different models, hardware configurations, and serving parameters.

Vendor:
Deployment:
Spec Decoding: