Time-Domain System Equivalent logoTime-Domain System EquivalentLinear dynamics, solved faster.Discuss an evaluation
SDK Documentation

Runtime performance tuning

Measure and tune the selected Runtime workload and execution target.

TDSE exposes one performance decision point: the options passed to tdse_model_create. There is no supported sequence of backend, precision, thread, or accelerator setters after creation.

Start with the defaults

tdse_model_t *model = NULL;
tdse_status_t status = tdse_model_create(
    pack_data, pack_size, NULL, NULL, &model);

This selects CPU, FP64, host-resident step data, the qualified OpenBLAS path, and the host's existing provider thread configuration.

Choose resources once

tdse_model_create_options_t options = tdse_model_create_options_init();
options.device = TDSE_DEVICE_CPU;
options.precision = TDSE_PRECISION_FP32;
options.cpu_threads = configured_provider_threads;
options.memory_limit_bytes = model_memory_limit;

tdse_status_t status = tdse_model_create(
    pack_data, pack_size, &options, NULL, &model);

The model keeps this execution identity until release. Recreate the model to change device, precision, residency, or CPU resources.

CPU ownership

TDSE deliberately does not call process-wide OpenBLAS thread setters and does not rewrite host affinity. The host configures its workers first; create options tell TDSE what must be observed.

  • cpu_threads = 0 inherits the current provider setting.
  • A positive cpu_threads value is an exact contract, not a hint.
  • cpu_ids is an exact logical-CPU set and must have cpu_threads entries.
  • TDSE_CPU_MEMORY_LOCAL requests row-local placement for Runtime-owned matrix memory when the platform can verify it.
  • A mismatch fails closed during preparation instead of silently using a different thread count or CPU set.

The TDSE benchmark command is the recommended host for choosing and validating CPU placement. It configures OpenBLAS and affinity outside the timed step.

GPU residency

TDSE_DATA_HOST keeps dynamic inputs and outputs at the normal host API boundary. TDSE_DATA_DEVICE requires a caller-owned queue and the generic device step APIs. Both modes keep the H matrix resident on the GPU after preparation. Device transfer and queue synchronization policy belong to the host integration and must not be confused with prepared-step compute latency.

Preparation and hot-path discipline

tdse_model_execution_prepare(...) performs the selected execution setup before the first step:

  • H conversion, placement, and accelerator upload;
  • provider thread and CPU-set validation;
  • graph or device-kernel preparation.

The complete host order is:

  1. create the model, validate the pack, and optionally query memory preflight
  2. call tdse_model_execution_prepare(...)
  3. optionally prewarm the prepared compute path with trial queries followed by tdse_step_discard(...); no committed state is advanced
  4. run physically meaningful initialization/pre-event steps and commit the host solver's accepted primary values when the application requires them
  5. enter the measured or deadline-critical interval

There is no generic Runtime warmup function because TDSE cannot invent a physically valid primary history for the host circuit. Benchmark warmup is an untimed measurement procedure, not a required application lifecycle stage. When zero prehistory is the intended initial condition, application history priming is unnecessary.

For applications where first-step latency matters, use this startup policy:

  • CPU: after preparation, run 1024 non-committing production-path trials;
  • GPU: use the same GPU, data location, borrowed queue, and buffers as production, run 1024 non-committing trials and confirm queue completion;
  • omit the separate trials when a physically valid initialization interval already exercises the same path for at least 1024 steps;
  • keep timers out of this policy and finish queue-completion checks before entering the real-time loop.

The exact call sequences, error handling, device-resident example, and cache limitations are documented in Warmup, Computational Prewarm, and Physical Initialization.

The prepared step must not allocate, upload fixed H data, collect telemetry, serialize results, or compute a correctness reference.

Verify the result

tdse_model_execution_info_t info = tdse_model_execution_info_init();
if (tdse_model_get_execution_info(model, &info) != TDSE_STATUS_OK) {
    /* handle error */
}

For an explicit CPU contract, check the active thread count and verification state before entering the real-time interval. For an accelerator, confirm the selected device and data location. A query is observational and never changes execution.

Measurement guidance

Use tdse benchmark rather than application-specific loops for formal measurements. It applies the same prepared-step definition across devices, runs correctness checks outside timing, records measurement dispersion, and emits the canonical compact result used by the report.

For the current command and C++ measurement controls, default values, fixed-count rules, and formal GPU recommendation, see Recommended Benchmark Warmup Settings.