Time-Domain System Equivalent logoTime-Domain System EquivalentLinear dynamics, solved faster.Discuss an evaluation
BenchmarksRuntime Performance

Runtime Performance Report

TDSE Runtime Performance

One prepared-step measurement contract across measured CPU, GPU, and FPGA platforms, with complete matrices, numerical verification, and explicit timing boundaries.

Status
Correctness verified
Report date
September 23, 2026
Platform groups
25
Verified cases
1299/1299

Executive Summary

Every reported CPU, GPU, and FPGA matrix case passed numerical verification.

The report contains 1299 formal measurements across 25 platform groups and 28 execution configurations: 16 CPU, 8 GPU, and 1 FPGA groups. All results use prepared-step latency as the comparable headline metric; device-specific timing boundaries remain explicit.

1299/1299verified cases
25platform groups
30common CPU/GPU points
FP32 + FP64precision coverage
How to compare

Compare only the same np, nh, precision, and prepared-step metric. The canonical report identifies 30 semantically equivalent matrix points shared by CPU and GPU platforms.

Prepared-Step Latency

Representative points from every measured configuration.

Values are medians of retained repeat means. Lower is faster. The PDF and CSV contain each platform's complete measured matrix.

PlatformCompact FP32np 8 · nh 512Mid-size FP32np 32 · nh 1024Mid-size FP64np 32 · nh 1024Large FP64np 128 · nh 2048
Windows 11 · AMD Ryzen 7 9800X3D (8 CPU cores)CPU · x86-642.21 µs5.68 µs11.86 µs3,841.08 µs
AWS c8i.2xlarge · Intel Xeon 6975P-C (4 CPU cores)CPU · x86-641.90 µs10.70 µs75.19 µs3,811.16 µs
AWS c8i.4xlarge · Intel Xeon 6975P-C (8 CPU cores)CPU · x86-641.85 µs6.56 µs11.35 µs1,994.11 µs
AWS c8i.8xlarge · Intel Xeon 6975P-C (16 CPU cores)CPU · x86-641.94 µs7.78 µs12.06 µs1,014.00 µs
AWS c8i.metal-48xl · Intel Xeon 6975P-C (96 CPU cores)CPU · x86-642.77 µs13.29 µs15.90 µs452.07 µs
AWS c8a.2xlarge · AMD EPYC 9R45 (4 CPU cores)CPU · x86-640.89 µs11.59 µs29.26 µs4,562.18 µs
AWS c8a.2xlarge · AMD EPYC 9R45 (8 CPU cores)CPU · x86-640.82 µs4.32 µs14.83 µs4,701.03 µs
AWS c8a.4xlarge · AMD EPYC 9R45 (16 CPU cores)CPU · x86-642.07 µs6.48 µs11.32 µs2,129.87 µs
AWS c8a.metal-24xl · AMD EPYC 9R45 (96 CPU cores)CPU · x86-642.98 µs11.11 µs14.25 µs292.61 µs
AWS c7i.4xlarge · Intel Xeon Platinum 8488C (8 CPU cores)CPU · x86-642.39 µs7.26 µs13.26 µs2,328.40 µs
AWS c7a.2xlarge · AMD EPYC 9R14 (8 CPU cores)CPU · x86-641.44 µs7.04 µs16.63 µs6,928.73 µs
AWS c9g.2xlarge · aarch64 (8 CPU cores)CPU · Arm641.67 µs13.33 µs24.67 µs1,743.84 µs
AWS c9g.4xlarge · aarch64 (16 CPU cores)CPU · Arm641.62 µs11.82 µs21.87 µs2,567.15 µs
AWS c8g.2xlarge · aarch64 (8 CPU cores)CPU · Arm641.95 µs17.99 µs31.91 µs1,930.25 µs
AWS c8g.4xlarge · aarch64 (16 CPU cores)CPU · Arm642.02 µs13.67 µs25.55 µs1,438.69 µs
macOS 27.0 · Apple M5 (10 CPU cores)CPU · Arm642.17 µs5.07 µs10.11 µs6,073.46 µs
Windows 11 · NVIDIA GeForce RTX 5090 (1 GPU)GPU · x86-6420.35 µs20.25 µs22.07 µs185.11 µs
RunPod · NVIDIA GeForce RTX 5090 (1 GPU)GPU · x86-6414.26 µs14.46 µs17.28 µs181.86 µs
RunPod · NVIDIA RTX PRO 6000 Blackwell (1 GPU)GPU · x86-6419.79 µs20.03 µs23.44 µs202.38 µs
RunPod · NVIDIA A100-SXM4-80GB (1 GPU)GPU · x86-6427.01 µs27.66 µs27.80 µs177.19 µs
RunPod · NVIDIA H100 80GB HBM3 (1 GPU)GPU · x86-6420.75 µs21.52 µs21.53 µs106.55 µs
RunPod · NVIDIA H200 (1 GPU)GPU · x86-6419.52 µs20.90 µs20.24 µs79.20 µs
RunPod · NVIDIA B200 (1 GPU)GPU · x86-6421.08 µs23.85 µs21.94 µs65.47 µs
RunPod · NVIDIA RTX PRO 6000 Blackwell Workstation Edition (1 GPU)GPU · x86-6416.81 µs16.52 µs19.94 µs186.45 µs
RunPod · NVIDIA RTX PRO 6000 Blackwell Workstation Edition (2 GPUs)GPU · x86-6425.16 µs26.24 µs40.46 µs82.92 µs
RunPod · NVIDIA RTX PRO 6000 Blackwell Workstation Edition (3 GPUs)GPU · x86-6446.88 µs54.06 µs52.89 µs76.94 µs
RunPod · NVIDIA RTX PRO 6000 Blackwell Workstation Edition (4 GPUs)GPU · x86-6452.96 µs57.03 µs59.12 µs110.52 µs
AWS f2.6xlarge · AMD/Xilinx Virtex UltraScale+ VU47P (1 FPGA)FPGA · x86-64—12.49 µs——

Accelerator Results

Prepared-step timing stays explicit across GPU and FPGA paths.

Windows 11 · NVIDIA GeForce RTX 5090 (1 GPU)

20.25 µs FP32

Device-resident: FP32 20.25 µs; FP64 22.07 µs.

Host-resident: FP32 24.08 µs; FP64 26.96 µs. Includes per-step vector transfers; H stays on the GPU.

Numerical max across this configuration: abs 2.29e-13, rel 1.40e-5.

RunPod · NVIDIA GeForce RTX 5090 (1 GPU)

14.46 µs FP32

Device-resident: FP32 14.46 µs; FP64 17.28 µs.

Host-resident: FP32 18.05 µs; FP64 20.44 µs. Includes per-step vector transfers; H stays on the GPU.

Numerical max across this configuration: abs 2.29e-13, rel 1.40e-5.

RunPod · NVIDIA RTX PRO 6000 Blackwell (1 GPU)

20.03 µs FP32

Device-resident: FP32 20.03 µs; FP64 23.44 µs.

Host-resident: FP32 22.50 µs; FP64 26.05 µs. Includes per-step vector transfers; H stays on the GPU.

Numerical max across this configuration: abs 2.39e-13, rel 1.40e-5.

RunPod · NVIDIA A100-SXM4-80GB (1 GPU)

27.66 µs FP32

Device-resident: FP32 27.66 µs; FP64 27.80 µs.

Host-resident: FP32 38.17 µs; FP64 38.45 µs. Includes per-step vector transfers; H stays on the GPU.

Numerical max across this configuration: abs 3.09e-13, rel 1.45e-5.

RunPod · NVIDIA H100 80GB HBM3 (1 GPU)

21.52 µs FP32

Device-resident: FP32 21.52 µs; FP64 21.53 µs.

Host-resident: FP32 28.34 µs; FP64 27.02 µs. Includes per-step vector transfers; H stays on the GPU.

Numerical max across this configuration: abs 3.69e-13, rel 1.01e-5.

RunPod · NVIDIA H200 (1 GPU)

20.90 µs FP32

Device-resident: FP32 20.90 µs; FP64 20.24 µs.

Host-resident: FP32 25.54 µs; FP64 25.67 µs. Includes per-step vector transfers; H stays on the GPU.

Numerical max across this configuration: abs 3.69e-13, rel 1.01e-5.

RunPod · NVIDIA B200 (1 GPU)

23.85 µs FP32

Device-resident: FP32 23.85 µs; FP64 21.94 µs.

Host-resident: FP32 27.71 µs; FP64 27.74 µs. Includes per-step vector transfers; H stays on the GPU.

Numerical max across this configuration: abs 3.44e-13, rel 3.16e-5.

RunPod · NVIDIA RTX PRO 6000 Blackwell Workstation Edition (1 GPU)

16.52 µs FP32

Device-resident: FP32 16.52 µs; FP64 19.94 µs.

Host-resident: FP32 23.09 µs; FP64 25.90 µs. Includes per-step vector transfers; H stays on the GPU.

Numerical max across this configuration: abs 2.39e-13, rel 1.40e-5.

RunPod · NVIDIA RTX PRO 6000 Blackwell Workstation Edition (2 GPUs)

26.24 µs FP32

Device-resident: FP32 26.24 µs; FP64 40.46 µs.

Host-resident: FP32 31.63 µs; FP64 33.88 µs. Includes per-step vector transfers; H stays on the GPU.

Numerical max across this configuration: abs 4.43e-13, rel 6.74e-6.

RunPod · NVIDIA RTX PRO 6000 Blackwell Workstation Edition (3 GPUs)

54.06 µs FP32

Device-resident: FP32 54.06 µs; FP64 52.89 µs.

Host-resident: FP32 40.14 µs; FP64 40.48 µs. Includes per-step vector transfers; H stays on the GPU.

Numerical max across this configuration: abs 4.43e-13, rel 6.74e-6.

RunPod · NVIDIA RTX PRO 6000 Blackwell Workstation Edition (4 GPUs)

57.03 µs FP32

Device-resident: FP32 57.03 µs; FP64 59.12 µs.

Host-resident: FP32 48.36 µs; FP64 48.43 µs. Includes per-step vector transfers; H stays on the GPU.

Numerical max across this configuration: abs 3.86e-13, rel 6.74e-6.

AWS f2.6xlarge · AMD/Xilinx Virtex UltraScale+ VU47P (1 FPGA)

12.49 µs FP32

FP32 fixed-bitstream result. Host-observed step includes vector transfers.

Numerical max across this configuration: abs 5.69e-3, rel 2.48e-7.

Multi-GPU scaling

More work to share. Less time per step.

256 ports · 2,048 history samples. The same large operator, distributed across one to four GPUs.

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

FP32

Median step latency; speedup relative to this configuration's one-GPU result.
GPUsDevice-residentSpeedupHost-residentSpeedup
1347.65 µs1.00×409.41 µs1.00×
2212.43 µs1.64×243.76 µs1.68×
3164.47 µs2.11×197.75 µs2.07×
482.99 µs4.19×122.11 µs3.35×

FP64

Median step latency; speedup relative to this configuration's one-GPU result.
GPUsDevice-residentSpeedupHost-residentSpeedup
1679.73 µs1.00×740.80 µs1.00×
2394.53 µs1.72×432.29 µs1.71×
3290.11 µs2.34×325.46 µs2.28×
4231.63 µs2.93×272.87 µs2.71×

Scaling depends on workload size. Smaller operators can remain faster on one GPU. Device-resident and host-resident results are separate measurements, not an isolated transfer-cost measurement.

Verified Configurations

One result grammar across processor and accelerator platforms.

PlatformDeviceArchitectureComputeCasesMedian CV
Windows 11 · AMD Ryzen 7 9800X3D (8 CPU cores)Windows-11-10.0.26200-SP0; 11; AMD64 (little-endian)CPUx86-648 physical cores48/480.88%
AWS c8i.2xlarge · Intel Xeon 6975P-C (4 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; x86_64 (little-endian)CPUx86-644 physical cores48/480.13%
AWS c8i.4xlarge · Intel Xeon 6975P-C (8 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; x86_64 (little-endian)CPUx86-648 physical cores48/480.10%
AWS c8i.8xlarge · Intel Xeon 6975P-C (16 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; x86_64 (little-endian)CPUx86-6416 physical cores48/480.11%
AWS c8i.metal-48xl · Intel Xeon 6975P-C (96 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; x86_64 (little-endian)CPUx86-6496 physical cores48/480.07%
AWS c8a.2xlarge · AMD EPYC 9R45 (4 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; x86_64 (little-endian)CPUx86-644 physical cores48/480.17%
AWS c8a.2xlarge · AMD EPYC 9R45 (8 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; x86_64 (little-endian)CPUx86-648 physical cores48/480.14%
AWS c8a.4xlarge · AMD EPYC 9R45 (16 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; x86_64 (little-endian)CPUx86-6416 physical cores48/480.13%
AWS c8a.metal-24xl · AMD EPYC 9R45 (96 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; x86_64 (little-endian)CPUx86-6496 physical cores48/480.20%
AWS c7i.4xlarge · Intel Xeon Platinum 8488C (8 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; x86_64 (little-endian)CPUx86-648 physical cores48/480.50%
AWS c7a.2xlarge · AMD EPYC 9R14 (8 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; x86_64 (little-endian)CPUx86-648 physical cores48/480.23%
AWS c9g.2xlarge · aarch64 (8 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; aarch64 (little-endian)CPUArm648 physical cores48/480.05%
AWS c9g.4xlarge · aarch64 (16 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; aarch64 (little-endian)CPUArm6416 physical cores48/480.04%
AWS c8g.2xlarge · aarch64 (8 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; aarch64 (little-endian)CPUArm648 physical cores48/480.06%
AWS c8g.4xlarge · aarch64 (16 CPU cores)Ubuntu 24.04.4 LTS; 7.0.0-1012-aws; aarch64 (little-endian)CPUArm6416 physical cores48/480.06%
macOS 27.0 · Apple M5 (10 CPU cores)macOS 27.0; 27.0.0; arm64 (little-endian)CPUArm6410 physical cores48/480.28%
Windows 11 · NVIDIA GeForce RTX 5090 (1 GPU)windows; x86_64GPUx86-64170 SMs · 31.8 GiB per GPU48/480.51%
RunPod · NVIDIA GeForce RTX 5090 (1 GPU)linux; x86_64GPUx86-64170 SMs · 31.4 GiB per GPU48/480.04%
RunPod · NVIDIA RTX PRO 6000 Blackwell (1 GPU)linux; x86_64GPUx86-64188 SMs · 95.0 GiB per GPU48/480.12%
RunPod · NVIDIA A100-SXM4-80GB (1 GPU)linux; x86_64GPUx86-64108 SMs · 79.2 GiB per GPU48/480.53%
RunPod · NVIDIA H100 80GB HBM3 (1 GPU)linux; x86_64GPUx86-64132 SMs · 79.2 GiB per GPU48/480.04%
RunPod · NVIDIA H200 (1 GPU)linux; x86_64GPUx86-64132 SMs · 139.8 GiB per GPU48/480.04%
RunPod · NVIDIA B200 (1 GPU)linux; x86_64GPUx86-64148 SMs · 178.4 GiB per GPU48/480.08%
RunPod · NVIDIA RTX PRO 6000 Blackwell Workstation Edition (1 GPU)linux; x86_64GPUx86-64188 SMs · 95.0 GiB per GPU48/480.11%
RunPod · NVIDIA RTX PRO 6000 Blackwell Workstation Edition (2 GPUs)linux; x86_64GPUx86-64188 SMs · 95.0 GiB per GPU48/480.31%
RunPod · NVIDIA RTX PRO 6000 Blackwell Workstation Edition (3 GPUs)linux; x86_64GPUx86-64188 SMs · 95.0 GiB per GPU48/480.88%
RunPod · NVIDIA RTX PRO 6000 Blackwell Workstation Edition (4 GPUs)linux; x86_64GPUx86-64188 SMs · 95.0 GiB per GPU48/480.69%
AWS f2.6xlarge · AMD/Xilinx Virtex UltraScale+ VU47P (1 FPGA)Ubuntu 24.04.1 LTS; 6.8.0-1021-aws; x86_64FPGAx86-64AMD/Xilinx Virtex UltraScale+ VU47P · AWS F2 Small Shell3/31.28%

Measurement Method

Prepared steps, with the timing boundary made clear.

  1. Measure the native implementation.CPU, GPU, and FPGA results use their native execution paths. Compare like-for-like workloads, precision, and vector residency.
  2. Prepare outside the timed region.Model preparation and operator loading precede timing. GPU completion waits and required per-step host transfers remain inside their stated timing boundaries.
  3. Warm, measure, and repeat.Latency is the median of retained repeat means under the report's recorded sample-selection rules. CPU/GPU and FPGA use different sampling protocols.
  4. Verify every case.Only complete, uninstrumented, numerically passing formal matrices enter this report.

Downloads

The website and PDF use the same canonical report data.