BMG Hardware Peak and MFU Reference
This note documents the Intel Arc Pro B60 peak throughput assumptions used by the ARK/Sage benchmark notes, including the source links and formulas behind the FP16, BF16, INT8, and FP32 denominators.
INT8 XMX Peak
FP16/BF16 XMX Peak
FP32 Vector Peak
For the 75.6k non-causal attention benchmark, use 98.3 TFLOP/s as the FP16/BF16 XMX
denominator and 196.6 TOPS as the INT8 XMX denominator. The old shorthand
FP16 = INT8 / 2 is correct for XMX peak, but the comment should say that 196.6 is the INT8
dense XMX peak, not the FP16 peak.
Source Facts
| Fact | Value | Source |
|---|---|---|
| B60 GPU peak TOPS | 197 TOPS, INT8 dense XMX workload | Intel Arc Pro B60 datasheet |
| B60 FP32 throughput | Up to 12.28 TFLOP/s | Intel Arc Pro B60 datasheet |
| B60 compute resources | 20 Xe-cores, 160 XMX engines | Intel Arc Pro B-series overview |
| Per Xe-core matrix throughput | 4096 INT8 ops/cycle, 2048 FP16/BF16 ops/cycle | Intel Xe GPU architecture guide |
| Per Xe-core vector throughput | 256 FP32 ops/cycle | Intel Xe GPU architecture guide |
Formulas
The B60 spec implies a maximum graphics clock around 2.4 GHz because the official INT8 and FP32 peak numbers both solve to that frequency.
Using 2.4 GHz for simple benchmark denominators:
MFU Denominator
Model FLOP utilization is a ratio between the benchmark's counted attention FLOP/s and the device's theoretical peak for the datatype and execution unit used by that kernel.
| Benchmark path | Measured throughput | Peak denominator | MFU formula | MFU |
|---|---|---|---|---|
| ARK FP16 SDPA, BMG 75.6k non-causal | 83.02 TFLOP/s | 98.3 TFLOP/s FP16 XMX peak | 83.02 / 98.3 |
84.45% |
| ARK Sage v1 INT8, BMG 75.6k non-causal | 98.38 TFLOP/s | 196.6 TOPS INT8 XMX peak | 98.38 / 196.6 |
50.04% |
Do not compare FP16 MFU and INT8 MFU without checking the denominator. INT8 can have better latency and higher absolute achieved throughput while reporting lower MFU because its theoretical XMX peak is twice the FP16/BF16 XMX peak on this architecture.
Benchmark Interpretation
-
The standalone sycl-tla BKC result around
82.7 TFLOP/sis plausible: it is about82.7 / 98.3 = 84.1%of B60 FP16/BF16 XMX peak. -
The PyTorch XPU SDPA result around
55.9 TFLOP/sis about55.9 / 98.3 = 56.9%of B60 FP16/BF16 XMX peak, so the gap is a backend/kernel efficiency issue rather than a hardware peak mismatch. -
The Sage v1 INT8 result around
98.4 TOPSis about half of INT8 dense XMX peak, but it is already roughly equal to the full FP16/BF16 XMX theoretical peak. That is why raw TFLOP/s and MFU tell different stories.