Runtime Performance Benchmark
Runtime performance measurement, benchmark configuration, and interpretation.
TDSE has one public performance workflow for both CPU and GPU:
tdse benchmark --device cpu|gpu. There is no calibration stage, profile
file, separate GPU benchmark executable, or silent fallback between devices.
Quick Start
Run the default CPU matrix or the bounded local GPU regression:
tdse benchmark --device cpu --output tdse_cpu_matrix.json --json-out -
tdse benchmark --device gpu --output tdse_gpu_matrix.json --json-out -
Run one shape:
tdse benchmark --device cpu --np 32 --nh 1024 --dtype 64 \
--output np32_nh1024_fp64.json --json-out -
Benchmark a real Runtime Pack:
tdse benchmark --device cpu --pack model.pack --dtype 64 \
--output model_matrix.json --json-out -
The default CPU matrix is the Cartesian product of:
np:1, 2, 4, 8, 16, 32, 64, 128nh:512, 1024, 2048- precision: fp32 and fp64
The default GPU local-regression profile is the following eight explicit
precision cases:
(np, nh):(8,512),(32,1024),(64,2048),(128,1024)- precision: fp32 and fp64 for every pair
It is diagnostic evidence and uses 2048 measured steps, 64 untimed warmup steps, five recorded repeats, and one settling repeat per integration mode. The 48-row GPU formal matrix remains available only through an explicit qualification request:
tdse benchmark --device gpu \
--device-ids 0,2,5,7 \
--matrix-profile formal --evidence-class formal \
--output tdse_gpu_formal_matrix.json --json-out -
CPU and GPU rows measure the same prepared TDSE step. GPU rows report
the GPU-resident prepared-step headline, a separately labelled end-to-end
step, and CPU-resident integration evidence. Both integration modes keep the
H matrix resident in device memory. --device gpu rejects Runtime Packs and
CPU-only resource or tail options.
On Windows, the GPU command temporarily pins its process to both SMT threads
of one complete physical CPU core selected from the process-visible topology.
GPU-resident and CPU-resident integration therefore use the same host
scheduling boundary. The command does not raise process priority or require a
non-default power plan, and restores the original process affinity on exit.
The artifact records the applied mask, logical-processor count, unchanged
priority policy, and observed priority class under
environment.host_execution_control. This is a profiler host control only;
the Runtime library never changes a customer process's affinity. The
standalone CPU benchmark retains its existing topology and priority policy.
--device-ids is the only multi-GPU selector. Its ordered list is passed
directly to model creation, with the first device as root. Accelerator
evidence records that exact list once, plus each selected device's contiguous
row range and resident-byte footprint for the matrix's largest-residency
case, the required directed root/worker P2P links, and committed-epoch
consistency. Unselected visible GPUs are not used.
CPU Execution Policy
The benchmark chooses one deterministic plan from the process-visible CPU topology before timing:
- It keeps one logical CPU representative for every visible physical core; SMT siblings are not counted as additional compute cores.
- A locality domain is the intersection of an OS NUMA memory node and the highest-level shared data/unified cache domain visible to the OS. Package and die identity are fallback evidence when LLC sharing cannot be read.
- Output rows are split into contiguous complete-row tasks. Tasks are grouped across active locality domains so their row counts are as even as possible.
- The OpenBLAS configured thread count is the number of selected physical
cores. OpenBLAS activates at most one task per output row, so
np=6on a 96-core host configures 96 threads but executes six complete-row tasks. - The approved OpenBLAS provider disables split-K GEMV. A matrix output row is therefore computed by exactly one task and is never shared by workers in different locality domains.
If the number of output rows is smaller than the number of locality domains, only as many domains as rows receive active tasks. For example, six rows and three domains produce three contiguous two-row groups.
On Linux and Windows, the benchmark programs and reads back each OpenBLAS
worker's exact CPU affinity before timing. The result records CPU, NUMA, and
LLC-domain ownership for every row task. The Windows provider extension uses
processor-group-aware worker masks, so machines with more than 64 logical
processors do not collapse into one processor group. Linux additionally binds
and verifies matrix pages by row/NUMA group before timing. Windows records
windows_host_default_memory_exact_row_workers: worker ownership is exact,
while page placement remains under the host OS.
Timing Contract
Model creation, pack conversion, history priming, topology discovery, worker affinity, memory placement, correctness replay, and result serialization are outside the measured step interval. The timed interval contains prepared TDSE steps only.
Profiler warmup is a measurement control. It runs untimed prepared steps to remove first-use effects from the reported samples, using the benchmark's deterministic workload and state schedule. It is not a customer Runtime API, and it must not be confused with the physically meaningful initialization history supplied by a circuit host. Customer applications use the same recommended count through non-committing production-path trials, as described in Step Execution.
Each row runs one repeat batch. The benchmark never reruns a row to obtain a preferred variability label. It reports the observed repeat means directly, together with:
- mean and median ns/step
- population standard deviation
- coefficient of variation
- median absolute deviation
- first-half versus last-half relative drift
- min, max, and p95 repeat mean
The measured-step budget is adaptive and bounded. Warmup is a fixed count. All controls are public and use the same names in CLI and C++:
| CLI | C++ field | Meaning |
|---|---|---|
--min-steps | minSteps | minimum measured steps per repeat |
--max-steps | maxSteps | maximum measured steps per repeat |
--target-repeat-ms | targetRepeatMs | target duration per repeat |
--warmup-steps | warmupSteps | exact untimed warmup calls per repeat |
--repeats | repeats | repeat count |
The CPU default and explicit GPU formal profile share one contract: 64 minimum and 1,000,000 maximum measured steps, an approximately 250 ms target per repeat, exactly 64 untimed warmup calls, and seven repeats. The untimed pilot chooses the measured-step count once for each row; all seven repeats then use that same count and schedule. The default GPU local regression uses its fixed, smaller budget described above.
For an adaptive GPU row, every schedule cycle is a complete sequential ring traversal. Both GPU integration modes are conditioned outside measurement for one repeat-target duration. GPU-resident calibration then uses a settling pilot that reaches the repeat target and an independent estimator pilot of the same step count. Route assertions are checked once per batch, outside the reported step interval. These GPU state controls are separate from the exact 64 warmup calls executed before each reported repeat.
Recommended Benchmark Warmup Settings
The formal product rule is intentionally simple: CPU and explicit GPU formal runs both execute exactly 64 untimed warmup calls per row and repeat. Managed qualification runs use that value. The bounded default GPU regression instead owns its fixed 64-step warmup and does not accept measurement overrides.
Use the formal defaults by omitting all measurement options:
tdse benchmark --device cpu --output cpu-matrix.json
tdse benchmark --device gpu --matrix-profile formal \
--evidence-class formal --output gpu-formal-matrix.json
Override the count only for a custom diagnostic experiment based on the formal shape profile:
tdse benchmark --device gpu \
--matrix-profile formal --evidence-class diagnostic \
--warmup-steps 64 \
--output gpu-matrix.json
Each repeat resets and primes its disposable benchmark state, performs the
requested warmup outside timing, and then starts the measured schedule. Every
row records actual_warmup_steps, which must equal the requested
warmup_steps. There is no warmup duration target, minimum/maximum pair, or
warmup cap flag.
The measured-step count remains adaptive: untimed pilot work estimates
prepared-step cost, then minSteps, maxSteps, and targetRepeatMs determine
the timed count. The ordinary target is 250 ms. This does not alter the fixed
warmup rule.
The C++ runPerformanceMatrix(...) default argument and a default-constructed
PerformanceMeasurementOptions both use this same product contract. Override
only the fields the application deliberately intends to change.
For a real pack with a finite IR sequence, TDSE uses the same requested fixed warmup count. Longer measured batches are composed from identical reset-and-primed IR windows; resets, priming, and window changes remain outside every timed interval. This prevents a legal finite-duration pack from running beyond its IR horizon without changing Runtime IR semantics.
CPU Host Resource Boundary
The default is all process-visible physical cores. A benchmark host may narrow that boundary:
tdse benchmark --thread-cap 8 --output matrix.json
tdse benchmark --cpu-set 2,5,7 --output matrix.json
tdse benchmark --numa-node 0 --output matrix.json # Linux
--cpu-set accepts physical-core representatives only; SMT siblings and
duplicates are rejected. --numa-node is Linux-only. These are benchmark-host
controls: the ordinary TDSE Runtime does not change the embedding process's
OpenBLAS thread count, affinity, environment, or NUMA policy.
The benchmark snapshots process-global OpenBLAS and host-affinity state, applies its plan before model preparation, and restores the snapshot after the row. A failed restore makes the result unavailable.
Automation entry point
The tdse benchmark command is the sole product entry point for local and
remote performance runs. Repository profiler libraries and backend catalogs
are implementation details of that command and are not customer SDK headers.
Result Contract
The matrix JSON is intentionally compact: one row contains scalar statistics and short evidence arrays, not simulation history or coefficient dumps. Key fields include:
- hardware and software fingerprints;
- correctness receipts and repeat-level timing statistics;
- CPU topology, provider, thread, locality, and row-ownership evidence for CPU rows;
- CUDA device, H-matrix residency, prepared-step, end-to-end, and both integration-mode summaries for GPU rows.
The CLI's --json-out document is a command envelope. --output is the matrix
artifact. Archive both when an audit trail needs command status and numeric
evidence.
Failure Policy
The formal CPU benchmark fails closed when:
- the approved OpenBLAS provider is unavailable;
- OpenBLAS is not
DYNAMIC_ARCHor a forcedOPENBLAS_CORETYPEis present; - the requested CPU boundary is invalid;
- the configured provider thread count cannot be applied or restored;
- the row task plan disagrees with the execution geometry;
- correctness, placement, or the optional realtime-tail check fails.
There is no fallback to a hand-written CPU backend and no qualitative variability failure. Variability remains numeric evidence for the reader.
The GPU benchmark likewise fails closed when CUDA support, the CUDA device, the requested product route, H-matrix residency, or correctness verification is unavailable. It never substitutes CPU execution for a requested GPU run.
