Time-Domain System Equivalent logoTime-Domain System EquivalentLinear dynamics, solved faster.Discuss an evaluation
SDK Documentation

Runtime Performance Benchmark

Runtime performance measurement, benchmark configuration, and interpretation.

TDSE has one public performance workflow for both CPU and GPU: tdse benchmark --device cpu|gpu. There is no calibration stage, profile file, separate GPU benchmark executable, or silent fallback between devices.

Quick Start

Run the default CPU matrix or the bounded local GPU regression:

tdse benchmark --device cpu --output tdse_cpu_matrix.json --json-out -
tdse benchmark --device gpu --output tdse_gpu_matrix.json --json-out -

Run one shape:

tdse benchmark --device cpu --np 32 --nh 1024 --dtype 64 \
  --output np32_nh1024_fp64.json --json-out -

Benchmark a real Runtime Pack:

tdse benchmark --device cpu --pack model.pack --dtype 64 \
  --output model_matrix.json --json-out -

The default CPU matrix is the Cartesian product of:

  • np: 1, 2, 4, 8, 16, 32, 64, 128
  • nh: 512, 1024, 2048
  • precision: fp32 and fp64

The default GPU local-regression profile is the following eight explicit precision cases:

  • (np, nh): (8,512), (32,1024), (64,2048), (128,1024)
  • precision: fp32 and fp64 for every pair

It is diagnostic evidence and uses 2048 measured steps, 64 untimed warmup steps, five recorded repeats, and one settling repeat per integration mode. The 48-row GPU formal matrix remains available only through an explicit qualification request:

tdse benchmark --device gpu \
  --device-ids 0,2,5,7 \
  --matrix-profile formal --evidence-class formal \
  --output tdse_gpu_formal_matrix.json --json-out -

CPU and GPU rows measure the same prepared TDSE step. GPU rows report the GPU-resident prepared-step headline, a separately labelled end-to-end step, and CPU-resident integration evidence. Both integration modes keep the H matrix resident in device memory. --device gpu rejects Runtime Packs and CPU-only resource or tail options.

On Windows, the GPU command temporarily pins its process to both SMT threads of one complete physical CPU core selected from the process-visible topology. GPU-resident and CPU-resident integration therefore use the same host scheduling boundary. The command does not raise process priority or require a non-default power plan, and restores the original process affinity on exit. The artifact records the applied mask, logical-processor count, unchanged priority policy, and observed priority class under environment.host_execution_control. This is a profiler host control only; the Runtime library never changes a customer process's affinity. The standalone CPU benchmark retains its existing topology and priority policy.

--device-ids is the only multi-GPU selector. Its ordered list is passed directly to model creation, with the first device as root. Accelerator evidence records that exact list once, plus each selected device's contiguous row range and resident-byte footprint for the matrix's largest-residency case, the required directed root/worker P2P links, and committed-epoch consistency. Unselected visible GPUs are not used.

CPU Execution Policy

The benchmark chooses one deterministic plan from the process-visible CPU topology before timing:

  1. It keeps one logical CPU representative for every visible physical core; SMT siblings are not counted as additional compute cores.
  2. A locality domain is the intersection of an OS NUMA memory node and the highest-level shared data/unified cache domain visible to the OS. Package and die identity are fallback evidence when LLC sharing cannot be read.
  3. Output rows are split into contiguous complete-row tasks. Tasks are grouped across active locality domains so their row counts are as even as possible.
  4. The OpenBLAS configured thread count is the number of selected physical cores. OpenBLAS activates at most one task per output row, so np=6 on a 96-core host configures 96 threads but executes six complete-row tasks.
  5. The approved OpenBLAS provider disables split-K GEMV. A matrix output row is therefore computed by exactly one task and is never shared by workers in different locality domains.

If the number of output rows is smaller than the number of locality domains, only as many domains as rows receive active tasks. For example, six rows and three domains produce three contiguous two-row groups.

On Linux and Windows, the benchmark programs and reads back each OpenBLAS worker's exact CPU affinity before timing. The result records CPU, NUMA, and LLC-domain ownership for every row task. The Windows provider extension uses processor-group-aware worker masks, so machines with more than 64 logical processors do not collapse into one processor group. Linux additionally binds and verifies matrix pages by row/NUMA group before timing. Windows records windows_host_default_memory_exact_row_workers: worker ownership is exact, while page placement remains under the host OS.

Timing Contract

Model creation, pack conversion, history priming, topology discovery, worker affinity, memory placement, correctness replay, and result serialization are outside the measured step interval. The timed interval contains prepared TDSE steps only.

Profiler warmup is a measurement control. It runs untimed prepared steps to remove first-use effects from the reported samples, using the benchmark's deterministic workload and state schedule. It is not a customer Runtime API, and it must not be confused with the physically meaningful initialization history supplied by a circuit host. Customer applications use the same recommended count through non-committing production-path trials, as described in Step Execution.

Each row runs one repeat batch. The benchmark never reruns a row to obtain a preferred variability label. It reports the observed repeat means directly, together with:

  • mean and median ns/step
  • population standard deviation
  • coefficient of variation
  • median absolute deviation
  • first-half versus last-half relative drift
  • min, max, and p95 repeat mean

The measured-step budget is adaptive and bounded. Warmup is a fixed count. All controls are public and use the same names in CLI and C++:

CLIC++ fieldMeaning
--min-stepsminStepsminimum measured steps per repeat
--max-stepsmaxStepsmaximum measured steps per repeat
--target-repeat-mstargetRepeatMstarget duration per repeat
--warmup-stepswarmupStepsexact untimed warmup calls per repeat
--repeatsrepeatsrepeat count

The CPU default and explicit GPU formal profile share one contract: 64 minimum and 1,000,000 maximum measured steps, an approximately 250 ms target per repeat, exactly 64 untimed warmup calls, and seven repeats. The untimed pilot chooses the measured-step count once for each row; all seven repeats then use that same count and schedule. The default GPU local regression uses its fixed, smaller budget described above.

For an adaptive GPU row, every schedule cycle is a complete sequential ring traversal. Both GPU integration modes are conditioned outside measurement for one repeat-target duration. GPU-resident calibration then uses a settling pilot that reaches the repeat target and an independent estimator pilot of the same step count. Route assertions are checked once per batch, outside the reported step interval. These GPU state controls are separate from the exact 64 warmup calls executed before each reported repeat.

The formal product rule is intentionally simple: CPU and explicit GPU formal runs both execute exactly 64 untimed warmup calls per row and repeat. Managed qualification runs use that value. The bounded default GPU regression instead owns its fixed 64-step warmup and does not accept measurement overrides.

Use the formal defaults by omitting all measurement options:

tdse benchmark --device cpu --output cpu-matrix.json
tdse benchmark --device gpu --matrix-profile formal \
  --evidence-class formal --output gpu-formal-matrix.json

Override the count only for a custom diagnostic experiment based on the formal shape profile:

tdse benchmark --device gpu \
  --matrix-profile formal --evidence-class diagnostic \
  --warmup-steps 64 \
  --output gpu-matrix.json

Each repeat resets and primes its disposable benchmark state, performs the requested warmup outside timing, and then starts the measured schedule. Every row records actual_warmup_steps, which must equal the requested warmup_steps. There is no warmup duration target, minimum/maximum pair, or warmup cap flag.

The measured-step count remains adaptive: untimed pilot work estimates prepared-step cost, then minSteps, maxSteps, and targetRepeatMs determine the timed count. The ordinary target is 250 ms. This does not alter the fixed warmup rule.

The C++ runPerformanceMatrix(...) default argument and a default-constructed PerformanceMeasurementOptions both use this same product contract. Override only the fields the application deliberately intends to change.

For a real pack with a finite IR sequence, TDSE uses the same requested fixed warmup count. Longer measured batches are composed from identical reset-and-primed IR windows; resets, priming, and window changes remain outside every timed interval. This prevents a legal finite-duration pack from running beyond its IR horizon without changing Runtime IR semantics.

CPU Host Resource Boundary

The default is all process-visible physical cores. A benchmark host may narrow that boundary:

tdse benchmark --thread-cap 8 --output matrix.json
tdse benchmark --cpu-set 2,5,7 --output matrix.json
tdse benchmark --numa-node 0 --output matrix.json   # Linux

--cpu-set accepts physical-core representatives only; SMT siblings and duplicates are rejected. --numa-node is Linux-only. These are benchmark-host controls: the ordinary TDSE Runtime does not change the embedding process's OpenBLAS thread count, affinity, environment, or NUMA policy.

The benchmark snapshots process-global OpenBLAS and host-affinity state, applies its plan before model preparation, and restores the snapshot after the row. A failed restore makes the result unavailable.

Automation entry point

The tdse benchmark command is the sole product entry point for local and remote performance runs. Repository profiler libraries and backend catalogs are implementation details of that command and are not customer SDK headers.

Result Contract

The matrix JSON is intentionally compact: one row contains scalar statistics and short evidence arrays, not simulation history or coefficient dumps. Key fields include:

  • hardware and software fingerprints;
  • correctness receipts and repeat-level timing statistics;
  • CPU topology, provider, thread, locality, and row-ownership evidence for CPU rows;
  • CUDA device, H-matrix residency, prepared-step, end-to-end, and both integration-mode summaries for GPU rows.

The CLI's --json-out document is a command envelope. --output is the matrix artifact. Archive both when an audit trail needs command status and numeric evidence.

Failure Policy

The formal CPU benchmark fails closed when:

  • the approved OpenBLAS provider is unavailable;
  • OpenBLAS is not DYNAMIC_ARCH or a forced OPENBLAS_CORETYPE is present;
  • the requested CPU boundary is invalid;
  • the configured provider thread count cannot be applied or restored;
  • the row task plan disagrees with the execution geometry;
  • correctness, placement, or the optional realtime-tail check fails.

There is no fallback to a hand-written CPU backend and no qualitative variability failure. Variability remains numeric evidence for the reader.

The GPU benchmark likewise fails closed when CUDA support, the CUDA device, the requested product route, H-matrix residency, or correctness verification is unavailable. It never substitutes CPU execution for a requested GPU run.