Time-Domain System Equivalent logoTime-Domain System EquivalentLinear dynamics, solved faster.Discuss an evaluation
SDK Documentation

Execution devices and performance

Backend selection, runtime plans, CUDA, precision, thread control, and benchmark guidance.

TDSE has one model-creation interface for every supported execution device:

tdse_model_create_options_t options = tdse_model_create_options_init();
tdse_model_t *model = NULL;

options.device = TDSE_DEVICE_CPU;
options.precision = TDSE_PRECISION_FP64;

tdse_status_t status = tdse_model_create(
    pack_data, pack_size, &options, NULL, &model);

Passing NULL instead of &options selects the same CPU, FP64, host-data defaults. Device choice, precision, CPU resources, data location, and memory limit are immutable properties of the model. To change them, release the model and create another one.

Create options

FieldMeaningDefault
deviceTDSE_DEVICE_CPU, TDSE_DEVICE_GPU, or TDSE_DEVICE_FPGACPU
precisionTDSE_PRECISION_FP32 or TDSE_PRECISION_FP64FP64
device_idaccelerator ordinal; -1 lets TDSE use the implementation default-1
device_id_count, device_idsordered GPU set for GPU execution; the first ID is the root GPUnone
data_locationlocation of dynamic step data at the API boundaryhost
device_queueborrowed caller-owned queue for device-resident executionNULL
cpu_threadsexact host-configured CPU provider thread count; 0 inherits0
cpu_id_count, cpu_idsexact logical CPU set expected from the hostinherited
cpu_memory_policydefault or row-local Runtime matrix placementdefault
memory_limit_bytesper-model long-lived Runtime memory limitunlimited
contextoptional advanced logging/guard contextNULL; Runtime creates an owned default

TDSE copies the CPU and GPU ID lists during creation. It borrows device_queue and an explicit context; the caller must keep borrowed objects alive until the model is released.

CPU execution

The CPU product path uses the qualified OpenBLAS provider. TDSE does not change process-wide OpenBLAS thread settings or operating-system affinity. The host owns those settings. When cpu_threads or cpu_ids is explicit, Runtime preparation verifies the requested contract and fails closed if the provider or worker placement does not match it.

const int32_t cpus[] = {2, 5, 7};
tdse_model_create_options_t options = tdse_model_create_options_init();
options.cpu_threads = 3;
options.cpu_id_count = 3;
options.cpu_ids = cpus;
options.cpu_memory_policy = TDSE_CPU_MEMORY_LOCAL;

tdse_status_t status = tdse_model_create(
    pack_data, pack_size, &options, NULL, &model);

Use explicit CPU IDs only when the host has already configured the BLAS workers to use that exact set. TDSE records and verifies the contract; it does not take ownership of the host scheduler.

GPU execution

GPU selection is explicit and fail-closed. A missing, unauthorized, or incompatible GPU provider never falls back to CPU.

Leaving device_id_count at zero preserves scalar device_id behavior and selects exactly one GPU. To select an ordered GPU set, leave device_id at -1 and provide unique, visible ordinals. The first ordinal is the root GPU; device-resident queues and external device pointers belong to it. TDSE copies the list before tdse_model_create returns. A provider that cannot satisfy the complete set fails closed.

const int32_t selected_gpus[] = {0, 2, 5, 7};
tdse_model_create_options_t options = tdse_model_create_options_init();
options.device = TDSE_DEVICE_GPU;
options.device_id_count = 4;
options.device_ids = selected_gpus;
options.data_location = TDSE_DATA_HOST;

Host-resident integration keeps the ordinary host step boundary:

tdse_model_create_options_t options = tdse_model_create_options_init();
options.device = TDSE_DEVICE_GPU;
options.precision = TDSE_PRECISION_FP32;
options.device_id = 0;
options.data_location = TDSE_DATA_HOST;

Device-resident integration supplies a caller-owned device queue:

tdse_model_create_options_t options = tdse_model_create_options_init();
options.device = TDSE_DEVICE_GPU;
options.precision = TDSE_PRECISION_FP32;
options.device_id = 0;
options.data_location = TDSE_DATA_DEVICE;
options.device_queue = caller_owned_queue;

For device-resident execution, use the vendor-neutral device step boundary: tdse_eval_history_device_typed, tdse_eval_history_device_view_typed, and tdse_commit_primary_device_typed. TDSE never destroys the borrowed queue. The caller is responsible for ordering producer, TDSE, and consumer work on that queue and must keep it alive until the model is released. Destroying a borrowed queue early violates the host ownership contract; TDSE cannot safely probe an already-invalid vendor handle.

The H matrix is prepared before timed stepping and remains device-resident for the prepared model. Fixed allocation, coefficient upload, graph preparation, telemetry, correctness reference work, and serialization are not part of the prepared-step measurement.

cuda_route is diagnostic evidence about the history input source, not another backend or a complete kernel name. mirrored means a rolling doubled ring is available; packed means a packed vector is the selected source. Graph execution can materialize mirrored input into its fixed packed address before calling cuBLAS, so read cuda_route together with hist_op_kind. The only formal history math provider is cuBLAS; the former ordinary custom GEMV fallback has no product route and is not built.

Before preparation, tdse_model_get_memory_preflight(...) reports the largest known per-device GPU payload for the exact ordered GPU set supplied at creation. The result is deliberately a lower bound: provider-private CUDA/cuBLAS memory is opaque and excluded. It is capacity evidence, not a device-selection heuristic; TDSE does not add, remove, or reorder GPUs based on the estimate.

FPGA execution

FPGA selection uses the same options:

tdse_model_create_options_t options = tdse_model_create_options_init();
options.device = TDSE_DEVICE_FPGA;
options.precision = TDSE_PRECISION_FP32;
options.device_id = 0;

Unsupported hardware, missing provider software, or incompatible runtime state causes creation or preparation to fail. TDSE does not substitute CPU or GPU execution.

Inspecting the selected resources

Use one read-only query for CPU, GPU, and FPGA models:

tdse_model_execution_info_t info = tdse_model_execution_info_init();
tdse_status_t status = tdse_model_get_execution_info(model, &info);

The result reports the selected device, precision, root device ordinal, ordered GPU set, data location, queue, requested and active CPU thread counts, copied CPU IDs, and verification state. The query does not reconfigure the model or the host.

Benchmarking

Use the tdse benchmark command for formal CPU or GPU prepared-step measurements. The benchmark product owns host thread configuration, affinity, warmup, repetition, correctness checks, and evidence capture. Application code should not reproduce private execution plans or diagnostic routes.

Here, warmup is part of the benchmark measurement protocol. Customer hosts prepare execution with tdse_model_execution_prepare(...), may optionally prewarm the prepared path with non-committing trial/discard calls, and create real history only from their own accepted circuit initialization steps.

The recommended starting point is 1024 non-committing calls on both CPU and GPU. Use the exact production device, precision, data location, queue, buffers, threads, and CPU set. See Step Execution for application call sequences and Runtime Performance Benchmark for benchmark settings.

Performance results are valid only for the recorded hardware, software, precision, matrix dimensions, and measurement settings. End-to-end integration latency must be reported separately from prepared-step latency.