Compare · GPUs & AI chips
Blackwell B200 SXM vs Gaudi 3
RANKED BY DENSE FP16/BF16 TENSOR TFLOPS PER CHIP
#3 VS #5
CHECKED 06 SEPT 2026
Side by side
Verdict
Blackwell B200 SXM ranks #3 of 20 on Dense FP16/BF16 tensor TFLOPS per chip (2.25 PFLOPS); Gaudi 3 ranks #5 (1.678 PFLOPS).
Rank 03 of 20 · leads
Dense FP16/BF16 tensor TFLOPS per chip
2.25 PFLOPS
MakerNVIDIA
Date2025
Verified04 Sept 2026
Evidence
NVIDIA HGX Platform page, HGX B200 row ("8x NVIDIA Blackwell SXM", footnote 4: "HGX B300 and HGX B200 shipping now"). FP16/BF16 Tensor Core is listed as 36 PFLOPS for the 8-GPU board. Footnote 2: dense is half the sparse spec. Dense board total is therefore 18 PFLOPS; per chip 18 / 8 = 2.25 PFLOPS. Total memory 1.4 TB => 180 GB HBM3E per SXM.Rank 05 of 20
Dense FP16/BF16 tensor TFLOPS per chip
1.678 PFLOPS
MakerIntel
Date2024
Verified04 Sept 2026
Evidence
Intel Gaudi 3 AI Accelerator white paper (Intel document 817486), product-comparison table: "BF16 MME TFLOPS 1678" for Gaudi 3. The same paper's MME-precision table lists BF16 and FP8 at 1678 TFLOPS and "FP16 (signed)" at 459 TFLOPS — so this row is the BF16 matrix figure. 128 GB HBM2e. Not an FP16=BF16 part in the NVIDIA/AMD sense; the difference is disclosed here rather than hidden.Try another pair
More Blackwell B200 SXM matchups
- Blackwell B200 SXM vs Instinct MI355X
- Blackwell B200 SXM vs Instinct MI350X
- Blackwell B200 SXM vs Blackwell Ultra B300 SXM
- Blackwell B200 SXM vs Instinct MI325X
- Blackwell B200 SXM vs Instinct MI300X
- Blackwell B200 SXM vs H200 SXM
- Blackwell B200 SXM vs H100 SXM
- Blackwell B200 SXM vs Instinct MI300A
- Blackwell B200 SXM vs TPU v6e (Trillium)
- Blackwell B200 SXM vs Trainium2
- Blackwell B200 SXM vs TPU v5p
- Blackwell B200 SXM vs Gaudi 2
Source Vendor product pages and official spec documents (AMD Instinct MI355X / MI350X / MI325X / MI300X / MI300A / MI250X / MI250 / MI210; NVIDIA…
Last checked 06 Sept 2026
