Runtime performance tuning
Measure and tune the selected Runtime workload and execution target.
TDSE exposes one performance decision point: the options passed to
tdse_model_create. There is no supported sequence of backend, precision,
thread, or accelerator setters after creation.
Start with the defaults
tdse_model_t *model = NULL;
tdse_status_t status = tdse_model_create(
pack_data, pack_size, NULL, NULL, &model);
This selects CPU, FP64, host-resident step data, the qualified OpenBLAS path, and the host's existing provider thread configuration.
Choose resources once
tdse_model_create_options_t options = tdse_model_create_options_init();
options.device = TDSE_DEVICE_CPU;
options.precision = TDSE_PRECISION_FP32;
options.cpu_threads = configured_provider_threads;
options.memory_limit_bytes = model_memory_limit;
tdse_status_t status = tdse_model_create(
pack_data, pack_size, &options, NULL, &model);
The model keeps this execution identity until release. Recreate the model to change device, precision, residency, or CPU resources.
CPU ownership
TDSE deliberately does not call process-wide OpenBLAS thread setters and does not rewrite host affinity. The host configures its workers first; create options tell TDSE what must be observed.
cpu_threads = 0inherits the current provider setting.- A positive
cpu_threadsvalue is an exact contract, not a hint. cpu_idsis an exact logical-CPU set and must havecpu_threadsentries.TDSE_CPU_MEMORY_LOCALrequests row-local placement for Runtime-owned matrix memory when the platform can verify it.- A mismatch fails closed during preparation instead of silently using a different thread count or CPU set.
The TDSE benchmark command is the recommended host for choosing and validating CPU placement. It configures OpenBLAS and affinity outside the timed step.
GPU residency
TDSE_DATA_HOST keeps dynamic inputs and outputs at the normal host API
boundary. TDSE_DATA_DEVICE requires a caller-owned queue and the generic
device step APIs. Both modes keep the H matrix resident on the GPU after
preparation. Device transfer and queue synchronization policy belong to the
host integration and must not be confused with prepared-step compute latency.
Preparation and hot-path discipline
tdse_model_execution_prepare(...) performs the selected execution setup
before the first step:
- H conversion, placement, and accelerator upload;
- provider thread and CPU-set validation;
- graph or device-kernel preparation.
The complete host order is:
- create the model, validate the pack, and optionally query memory preflight
- call
tdse_model_execution_prepare(...) - optionally prewarm the prepared compute path with trial queries followed by
tdse_step_discard(...); no committed state is advanced - run physically meaningful initialization/pre-event steps and commit the host solver's accepted primary values when the application requires them
- enter the measured or deadline-critical interval
There is no generic Runtime warmup function because TDSE cannot invent a physically valid primary history for the host circuit. Benchmark warmup is an untimed measurement procedure, not a required application lifecycle stage. When zero prehistory is the intended initial condition, application history priming is unnecessary.
For applications where first-step latency matters, use this startup policy:
- CPU: after preparation, run 1024 non-committing production-path trials;
- GPU: use the same GPU, data location, borrowed queue, and buffers as production, run 1024 non-committing trials and confirm queue completion;
- omit the separate trials when a physically valid initialization interval already exercises the same path for at least 1024 steps;
- keep timers out of this policy and finish queue-completion checks before entering the real-time loop.
The exact call sequences, error handling, device-resident example, and cache limitations are documented in Warmup, Computational Prewarm, and Physical Initialization.
The prepared step must not allocate, upload fixed H data, collect telemetry, serialize results, or compute a correctness reference.
Verify the result
tdse_model_execution_info_t info = tdse_model_execution_info_init();
if (tdse_model_get_execution_info(model, &info) != TDSE_STATUS_OK) {
/* handle error */
}
For an explicit CPU contract, check the active thread count and verification state before entering the real-time interval. For an accelerator, confirm the selected device and data location. A query is observational and never changes execution.
Measurement guidance
Use tdse benchmark rather than application-specific loops for
formal measurements. It applies the same prepared-step definition across
devices, runs correctness checks outside timing, records measurement
dispersion, and emits the canonical compact result used by the report.
For the current command and C++ measurement controls, default values, fixed-count rules, and formal GPU recommendation, see Recommended Benchmark Warmup Settings.
