Executive Summary
Every reported CPU, GPU, and FPGA matrix case passed numerical verification.
The report contains 1299 formal measurements across 25 platform groups and 28 execution configurations: 16 CPU, 8 GPU, and 1 FPGA groups. All results use prepared-step latency as the comparable headline metric; device-specific timing boundaries remain explicit.
Compare only the same np, nh, precision, and prepared-step metric. The canonical report identifies 30 semantically equivalent matrix points shared by CPU and GPU platforms.
Prepared-Step Latency
Representative points from every measured configuration.
Values are medians of retained repeat means. Lower is faster. The PDF and CSV contain each platform's complete measured matrix.
| Platform | Compact FP32np 8 · nh 512 | Mid-size FP32np 32 · nh 1024 | Mid-size FP64np 32 · nh 1024 | Large FP64np 128 · nh 2048 |
|---|---|---|---|---|
| Windows 11 · AMD Ryzen 7 9800X3D (8 CPU cores)CPU · x86-64 | 2.21 µs | 5.68 µs | 11.86 µs | 3,841.08 µs |
| AWS c8i.2xlarge · Intel Xeon 6975P-C (4 CPU cores)CPU · x86-64 | 1.90 µs | 10.70 µs | 75.19 µs | 3,811.16 µs |
| AWS c8i.4xlarge · Intel Xeon 6975P-C (8 CPU cores)CPU · x86-64 | 1.85 µs | 6.56 µs | 11.35 µs | 1,994.11 µs |
| AWS c8i.8xlarge · Intel Xeon 6975P-C (16 CPU cores)CPU · x86-64 | 1.94 µs | 7.78 µs | 12.06 µs | 1,014.00 µs |
| AWS c8i.metal-48xl · Intel Xeon 6975P-C (96 CPU cores)CPU · x86-64 | 2.77 µs | 13.29 µs | 15.90 µs | 452.07 µs |
| AWS c8a.2xlarge · AMD EPYC 9R45 (4 CPU cores)CPU · x86-64 | 0.89 µs | 11.59 µs | 29.26 µs | 4,562.18 µs |
| AWS c8a.2xlarge · AMD EPYC 9R45 (8 CPU cores)CPU · x86-64 | 0.82 µs | 4.32 µs | 14.83 µs | 4,701.03 µs |
| AWS c8a.4xlarge · AMD EPYC 9R45 (16 CPU cores)CPU · x86-64 | 2.07 µs | 6.48 µs | 11.32 µs | 2,129.87 µs |
| AWS c8a.metal-24xl · AMD EPYC 9R45 (96 CPU cores)CPU · x86-64 | 2.98 µs | 11.11 µs | 14.25 µs | 292.61 µs |
| AWS c7i.4xlarge · Intel Xeon Platinum 8488C (8 CPU cores)CPU · x86-64 | 2.39 µs | 7.26 µs | 13.26 µs | 2,328.40 µs |
| AWS c7a.2xlarge · AMD EPYC 9R14 (8 CPU cores)CPU · x86-64 | 1.44 µs | 7.04 µs | 16.63 µs | 6,928.73 µs |
| AWS c9g.2xlarge · aarch64 (8 CPU cores)CPU · Arm64 | 1.67 µs | 13.33 µs | 24.67 µs | 1,743.84 µs |
| AWS c9g.4xlarge · aarch64 (16 CPU cores)CPU · Arm64 | 1.62 µs | 11.82 µs | 21.87 µs | 2,567.15 µs |
| AWS c8g.2xlarge · aarch64 (8 CPU cores)CPU · Arm64 | 1.95 µs | 17.99 µs | 31.91 µs | 1,930.25 µs |
| AWS c8g.4xlarge · aarch64 (16 CPU cores)CPU · Arm64 | 2.02 µs | 13.67 µs | 25.55 µs | 1,438.69 µs |
| macOS 27.0 · Apple M5 (10 CPU cores)CPU · Arm64 | 2.17 µs | 5.07 µs | 10.11 µs | 6,073.46 µs |
| Windows 11 · NVIDIA GeForce RTX 5090 (1 GPU)GPU · x86-64 | 20.35 µs | 20.25 µs | 22.07 µs | 185.11 µs |
| RunPod · NVIDIA GeForce RTX 5090 (1 GPU)GPU · x86-64 | 14.26 µs | 14.46 µs | 17.28 µs | 181.86 µs |
| RunPod · NVIDIA RTX PRO 6000 Blackwell (1 GPU)GPU · x86-64 | 19.79 µs | 20.03 µs | 23.44 µs | 202.38 µs |
| RunPod · NVIDIA A100-SXM4-80GB (1 GPU)GPU · x86-64 | 27.01 µs | 27.66 µs | 27.80 µs | 177.19 µs |
| RunPod · NVIDIA H100 80GB HBM3 (1 GPU)GPU · x86-64 | 20.75 µs | 21.52 µs | 21.53 µs | 106.55 µs |
| RunPod · NVIDIA H200 (1 GPU)GPU · x86-64 | 19.52 µs | 20.90 µs | 20.24 µs | 79.20 µs |
| RunPod · NVIDIA B200 (1 GPU)GPU · x86-64 | 21.08 µs | 23.85 µs | 21.94 µs | 65.47 µs |
| RunPod · NVIDIA RTX PRO 6000 Blackwell Workstation Edition (1 GPU)GPU · x86-64 | 16.81 µs | 16.52 µs | 19.94 µs | 186.45 µs |
| RunPod · NVIDIA RTX PRO 6000 Blackwell Workstation Edition (2 GPUs)GPU · x86-64 | 25.16 µs | 26.24 µs | 40.46 µs | 82.92 µs |
| RunPod · NVIDIA RTX PRO 6000 Blackwell Workstation Edition (3 GPUs)GPU · x86-64 | 46.88 µs | 54.06 µs | 52.89 µs | 76.94 µs |
| RunPod · NVIDIA RTX PRO 6000 Blackwell Workstation Edition (4 GPUs)GPU · x86-64 | 52.96 µs | 57.03 µs | 59.12 µs | 110.52 µs |
| AWS f2.6xlarge · AMD/Xilinx Virtex UltraScale+ VU47P (1 FPGA)FPGA · x86-64 | — | 12.49 µs | — | — |
Accelerator Results
Prepared-step timing stays explicit across GPU and FPGA paths.
20.25 µs FP32
Device-resident: FP32 20.25 µs; FP64 22.07 µs.
Host-resident: FP32 24.08 µs; FP64 26.96 µs. Includes per-step vector transfers; H stays on the GPU.
Numerical max across this configuration: abs 2.29e-13, rel 1.40e-5.
14.46 µs FP32
Device-resident: FP32 14.46 µs; FP64 17.28 µs.
Host-resident: FP32 18.05 µs; FP64 20.44 µs. Includes per-step vector transfers; H stays on the GPU.
Numerical max across this configuration: abs 2.29e-13, rel 1.40e-5.
20.03 µs FP32
Device-resident: FP32 20.03 µs; FP64 23.44 µs.
Host-resident: FP32 22.50 µs; FP64 26.05 µs. Includes per-step vector transfers; H stays on the GPU.
Numerical max across this configuration: abs 2.39e-13, rel 1.40e-5.
27.66 µs FP32
Device-resident: FP32 27.66 µs; FP64 27.80 µs.
Host-resident: FP32 38.17 µs; FP64 38.45 µs. Includes per-step vector transfers; H stays on the GPU.
Numerical max across this configuration: abs 3.09e-13, rel 1.45e-5.
21.52 µs FP32
Device-resident: FP32 21.52 µs; FP64 21.53 µs.
Host-resident: FP32 28.34 µs; FP64 27.02 µs. Includes per-step vector transfers; H stays on the GPU.
Numerical max across this configuration: abs 3.69e-13, rel 1.01e-5.
20.90 µs FP32
Device-resident: FP32 20.90 µs; FP64 20.24 µs.
Host-resident: FP32 25.54 µs; FP64 25.67 µs. Includes per-step vector transfers; H stays on the GPU.
Numerical max across this configuration: abs 3.69e-13, rel 1.01e-5.
23.85 µs FP32
Device-resident: FP32 23.85 µs; FP64 21.94 µs.
Host-resident: FP32 27.71 µs; FP64 27.74 µs. Includes per-step vector transfers; H stays on the GPU.
Numerical max across this configuration: abs 3.44e-13, rel 3.16e-5.
16.52 µs FP32
Device-resident: FP32 16.52 µs; FP64 19.94 µs.
Host-resident: FP32 23.09 µs; FP64 25.90 µs. Includes per-step vector transfers; H stays on the GPU.
Numerical max across this configuration: abs 2.39e-13, rel 1.40e-5.
26.24 µs FP32
Device-resident: FP32 26.24 µs; FP64 40.46 µs.
Host-resident: FP32 31.63 µs; FP64 33.88 µs. Includes per-step vector transfers; H stays on the GPU.
Numerical max across this configuration: abs 4.43e-13, rel 6.74e-6.
54.06 µs FP32
Device-resident: FP32 54.06 µs; FP64 52.89 µs.
Host-resident: FP32 40.14 µs; FP64 40.48 µs. Includes per-step vector transfers; H stays on the GPU.
Numerical max across this configuration: abs 4.43e-13, rel 6.74e-6.
57.03 µs FP32
Device-resident: FP32 57.03 µs; FP64 59.12 µs.
Host-resident: FP32 48.36 µs; FP64 48.43 µs. Includes per-step vector transfers; H stays on the GPU.
Numerical max across this configuration: abs 3.86e-13, rel 6.74e-6.
12.49 µs FP32
FP32 fixed-bitstream result. Host-observed step includes vector transfers.
Numerical max across this configuration: abs 5.69e-3, rel 2.48e-7.
Multi-GPU scaling
More work to share. Less time per step.
256 ports · 2,048 history samples. The same large operator, distributed across one to four GPUs.
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
FP32
| GPUs | Device-resident | Speedup | Host-resident | Speedup |
|---|---|---|---|---|
| 1 | 347.65 µs | 1.00× | 409.41 µs | 1.00× |
| 2 | 212.43 µs | 1.64× | 243.76 µs | 1.68× |
| 3 | 164.47 µs | 2.11× | 197.75 µs | 2.07× |
| 4 | 82.99 µs | 4.19× | 122.11 µs | 3.35× |
FP64
| GPUs | Device-resident | Speedup | Host-resident | Speedup |
|---|---|---|---|---|
| 1 | 679.73 µs | 1.00× | 740.80 µs | 1.00× |
| 2 | 394.53 µs | 1.72× | 432.29 µs | 1.71× |
| 3 | 290.11 µs | 2.34× | 325.46 µs | 2.28× |
| 4 | 231.63 µs | 2.93× | 272.87 µs | 2.71× |
Scaling depends on workload size. Smaller operators can remain faster on one GPU. Device-resident and host-resident results are separate measurements, not an isolated transfer-cost measurement.
Verified Configurations
One result grammar across processor and accelerator platforms.
| Platform | Device | Architecture | Compute | Cases | Median CV |
|---|---|---|---|---|---|
| Windows 11 · AMD Ryzen 7 9800X3D (8 CPU cores)Windows-11-10.0.26200-SP0; 11; AMD64 (little-endian) | CPU | x86-64 | 8 physical cores | 48/48 | 0.88% |
| AWS c8i.2xlarge · Intel Xeon 6975P-C (4 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; x86_64 (little-endian) | CPU | x86-64 | 4 physical cores | 48/48 | 0.13% |
| AWS c8i.4xlarge · Intel Xeon 6975P-C (8 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; x86_64 (little-endian) | CPU | x86-64 | 8 physical cores | 48/48 | 0.10% |
| AWS c8i.8xlarge · Intel Xeon 6975P-C (16 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; x86_64 (little-endian) | CPU | x86-64 | 16 physical cores | 48/48 | 0.11% |
| AWS c8i.metal-48xl · Intel Xeon 6975P-C (96 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; x86_64 (little-endian) | CPU | x86-64 | 96 physical cores | 48/48 | 0.07% |
| AWS c8a.2xlarge · AMD EPYC 9R45 (4 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; x86_64 (little-endian) | CPU | x86-64 | 4 physical cores | 48/48 | 0.17% |
| AWS c8a.2xlarge · AMD EPYC 9R45 (8 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; x86_64 (little-endian) | CPU | x86-64 | 8 physical cores | 48/48 | 0.14% |
| AWS c8a.4xlarge · AMD EPYC 9R45 (16 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; x86_64 (little-endian) | CPU | x86-64 | 16 physical cores | 48/48 | 0.13% |
| AWS c8a.metal-24xl · AMD EPYC 9R45 (96 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; x86_64 (little-endian) | CPU | x86-64 | 96 physical cores | 48/48 | 0.20% |
| AWS c7i.4xlarge · Intel Xeon Platinum 8488C (8 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; x86_64 (little-endian) | CPU | x86-64 | 8 physical cores | 48/48 | 0.50% |
| AWS c7a.2xlarge · AMD EPYC 9R14 (8 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; x86_64 (little-endian) | CPU | x86-64 | 8 physical cores | 48/48 | 0.23% |
| AWS c9g.2xlarge · aarch64 (8 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; aarch64 (little-endian) | CPU | Arm64 | 8 physical cores | 48/48 | 0.05% |
| AWS c9g.4xlarge · aarch64 (16 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; aarch64 (little-endian) | CPU | Arm64 | 16 physical cores | 48/48 | 0.04% |
| AWS c8g.2xlarge · aarch64 (8 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; aarch64 (little-endian) | CPU | Arm64 | 8 physical cores | 48/48 | 0.06% |
| AWS c8g.4xlarge · aarch64 (16 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; aarch64 (little-endian) | CPU | Arm64 | 16 physical cores | 48/48 | 0.06% |
| macOS 27.0 · Apple M5 (10 CPU cores)macOS 27.0; 27.0.0; arm64 (little-endian) | CPU | Arm64 | 10 physical cores | 48/48 | 0.28% |
| Windows 11 · NVIDIA GeForce RTX 5090 (1 GPU)windows; x86_64 | GPU | x86-64 | 170 SMs · 31.8 GiB per GPU | 48/48 | 0.51% |
| RunPod · NVIDIA GeForce RTX 5090 (1 GPU)linux; x86_64 | GPU | x86-64 | 170 SMs · 31.4 GiB per GPU | 48/48 | 0.04% |
| RunPod · NVIDIA RTX PRO 6000 Blackwell (1 GPU)linux; x86_64 | GPU | x86-64 | 188 SMs · 95.0 GiB per GPU | 48/48 | 0.12% |
| RunPod · NVIDIA A100-SXM4-80GB (1 GPU)linux; x86_64 | GPU | x86-64 | 108 SMs · 79.2 GiB per GPU | 48/48 | 0.53% |
| RunPod · NVIDIA H100 80GB HBM3 (1 GPU)linux; x86_64 | GPU | x86-64 | 132 SMs · 79.2 GiB per GPU | 48/48 | 0.04% |
| RunPod · NVIDIA H200 (1 GPU)linux; x86_64 | GPU | x86-64 | 132 SMs · 139.8 GiB per GPU | 48/48 | 0.04% |
| RunPod · NVIDIA B200 (1 GPU)linux; x86_64 | GPU | x86-64 | 148 SMs · 178.4 GiB per GPU | 48/48 | 0.08% |
| RunPod · NVIDIA RTX PRO 6000 Blackwell Workstation Edition (1 GPU)linux; x86_64 | GPU | x86-64 | 188 SMs · 95.0 GiB per GPU | 48/48 | 0.11% |
| RunPod · NVIDIA RTX PRO 6000 Blackwell Workstation Edition (2 GPUs)linux; x86_64 | GPU | x86-64 | 188 SMs · 95.0 GiB per GPU | 48/48 | 0.31% |
| RunPod · NVIDIA RTX PRO 6000 Blackwell Workstation Edition (3 GPUs)linux; x86_64 | GPU | x86-64 | 188 SMs · 95.0 GiB per GPU | 48/48 | 0.88% |
| RunPod · NVIDIA RTX PRO 6000 Blackwell Workstation Edition (4 GPUs)linux; x86_64 | GPU | x86-64 | 188 SMs · 95.0 GiB per GPU | 48/48 | 0.69% |
| AWS f2.6xlarge · AMD/Xilinx Virtex UltraScale+ VU47P (1 FPGA)Ubuntu 24.04.1 LTS; 6.8.0-1021-aws; x86_64 | FPGA | x86-64 | AMD/Xilinx Virtex UltraScale+ VU47P · AWS F2 Small Shell | 3/3 | 1.28% |
Measurement Method
Prepared steps, with the timing boundary made clear.
- Measure the native implementation.CPU, GPU, and FPGA results use their native execution paths. Compare like-for-like workloads, precision, and vector residency.
- Prepare outside the timed region.Model preparation and operator loading precede timing. GPU completion waits and required per-step host transfers remain inside their stated timing boundaries.
- Warm, measure, and repeat.Latency is the median of retained repeat means under the report's recorded sample-selection rules. CPU/GPU and FPGA use different sampling protocols.
- Verify every case.Only complete, uninstrumented, numerically passing formal matrices enter this report.
Downloads
