Execution devices and performance
Backend selection, runtime plans, CUDA, precision, thread control, and benchmark guidance.
TDSE has one model-creation interface for every supported execution device:
tdse_model_create_options_t options = tdse_model_create_options_init();
tdse_model_t *model = NULL;
options.device = TDSE_DEVICE_CPU;
options.precision = TDSE_PRECISION_FP64;
tdse_status_t status = tdse_model_create(
pack_data, pack_size, &options, NULL, &model);
Passing NULL instead of &options selects the same CPU, FP64, host-data
defaults. Device choice, precision, CPU resources, data location, and memory
limit are immutable properties of the model. To change them, release the model
and create another one.
Create options
| Field | Meaning | Default |
|---|---|---|
device | TDSE_DEVICE_CPU, TDSE_DEVICE_GPU, or TDSE_DEVICE_FPGA | CPU |
precision | TDSE_PRECISION_FP32 or TDSE_PRECISION_FP64 | FP64 |
device_id | accelerator ordinal; -1 lets TDSE use the implementation default | -1 |
device_id_count, device_ids | ordered GPU set for GPU execution; the first ID is the root GPU | none |
data_location | location of dynamic step data at the API boundary | host |
device_queue | borrowed caller-owned queue for device-resident execution | NULL |
cpu_threads | exact host-configured CPU provider thread count; 0 inherits | 0 |
cpu_id_count, cpu_ids | exact logical CPU set expected from the host | inherited |
cpu_memory_policy | default or row-local Runtime matrix placement | default |
memory_limit_bytes | per-model long-lived Runtime memory limit | unlimited |
context | optional advanced logging/guard context | NULL; Runtime creates an owned default |
TDSE copies the CPU and GPU ID lists during creation. It borrows device_queue
and an explicit context; the caller must keep borrowed objects alive until
the model is released.
CPU execution
The CPU product path uses the qualified OpenBLAS provider. TDSE does not change
process-wide OpenBLAS thread settings or operating-system affinity. The host
owns those settings. When cpu_threads or cpu_ids is explicit, Runtime
preparation verifies the requested contract and fails closed if the provider or
worker placement does not match it.
const int32_t cpus[] = {2, 5, 7};
tdse_model_create_options_t options = tdse_model_create_options_init();
options.cpu_threads = 3;
options.cpu_id_count = 3;
options.cpu_ids = cpus;
options.cpu_memory_policy = TDSE_CPU_MEMORY_LOCAL;
tdse_status_t status = tdse_model_create(
pack_data, pack_size, &options, NULL, &model);
Use explicit CPU IDs only when the host has already configured the BLAS workers to use that exact set. TDSE records and verifies the contract; it does not take ownership of the host scheduler.
GPU execution
GPU selection is explicit and fail-closed. A missing, unauthorized, or incompatible GPU provider never falls back to CPU.
Leaving device_id_count at zero preserves scalar device_id behavior and
selects exactly one GPU. To select an ordered GPU set, leave device_id at
-1 and provide unique, visible ordinals. The first ordinal is the root GPU;
device-resident queues and external device pointers belong to it. TDSE copies
the list before tdse_model_create returns. A provider that cannot satisfy the
complete set fails closed.
const int32_t selected_gpus[] = {0, 2, 5, 7};
tdse_model_create_options_t options = tdse_model_create_options_init();
options.device = TDSE_DEVICE_GPU;
options.device_id_count = 4;
options.device_ids = selected_gpus;
options.data_location = TDSE_DATA_HOST;
Host-resident integration keeps the ordinary host step boundary:
tdse_model_create_options_t options = tdse_model_create_options_init();
options.device = TDSE_DEVICE_GPU;
options.precision = TDSE_PRECISION_FP32;
options.device_id = 0;
options.data_location = TDSE_DATA_HOST;
Device-resident integration supplies a caller-owned device queue:
tdse_model_create_options_t options = tdse_model_create_options_init();
options.device = TDSE_DEVICE_GPU;
options.precision = TDSE_PRECISION_FP32;
options.device_id = 0;
options.data_location = TDSE_DATA_DEVICE;
options.device_queue = caller_owned_queue;
For device-resident execution, use the vendor-neutral device step boundary:
tdse_eval_history_device_typed,
tdse_eval_history_device_view_typed, and
tdse_commit_primary_device_typed. TDSE never destroys the borrowed queue.
The caller is responsible for ordering producer, TDSE, and consumer work on
that queue and must keep it alive until the model is released. Destroying a
borrowed queue early violates the host ownership contract; TDSE cannot safely
probe an already-invalid vendor handle.
The H matrix is prepared before timed stepping and remains device-resident for the prepared model. Fixed allocation, coefficient upload, graph preparation, telemetry, correctness reference work, and serialization are not part of the prepared-step measurement.
cuda_route is diagnostic evidence about the history input source, not another
backend or a complete kernel name. mirrored means a rolling doubled ring is
available; packed means a packed vector is the selected source. Graph execution
can materialize mirrored input into its fixed packed address before calling
cuBLAS, so read cuda_route together with hist_op_kind. The only formal history
math provider is cuBLAS; the former ordinary custom GEMV fallback has no product
route and is not built.
Before preparation, tdse_model_get_memory_preflight(...) reports the largest
known per-device GPU payload for the exact ordered GPU set supplied at creation.
The result is deliberately a lower bound: provider-private CUDA/cuBLAS memory is
opaque and excluded. It is capacity evidence, not a device-selection heuristic;
TDSE does not add, remove, or reorder GPUs based on the estimate.
FPGA execution
FPGA selection uses the same options:
tdse_model_create_options_t options = tdse_model_create_options_init();
options.device = TDSE_DEVICE_FPGA;
options.precision = TDSE_PRECISION_FP32;
options.device_id = 0;
Unsupported hardware, missing provider software, or incompatible runtime state causes creation or preparation to fail. TDSE does not substitute CPU or GPU execution.
Inspecting the selected resources
Use one read-only query for CPU, GPU, and FPGA models:
tdse_model_execution_info_t info = tdse_model_execution_info_init();
tdse_status_t status = tdse_model_get_execution_info(model, &info);
The result reports the selected device, precision, root device ordinal, ordered GPU set, data location, queue, requested and active CPU thread counts, copied CPU IDs, and verification state. The query does not reconfigure the model or the host.
Benchmarking
Use the tdse benchmark command for formal CPU or GPU prepared-step
measurements. The benchmark product owns host thread configuration, affinity,
warmup, repetition, correctness checks, and evidence capture. Application code
should not reproduce private execution plans or diagnostic routes.
Here, warmup is part of the benchmark measurement protocol. Customer hosts
prepare execution with tdse_model_execution_prepare(...), may optionally
prewarm the prepared path with non-committing trial/discard calls, and create
real history only from their own accepted circuit initialization steps.
The recommended starting point is 1024 non-committing calls on both CPU and GPU. Use the exact production device, precision, data location, queue, buffers, threads, and CPU set. See Step Execution for application call sequences and Runtime Performance Benchmark for benchmark settings.
Performance results are valid only for the recorded hardware, software, precision, matrix dimensions, and measurement settings. End-to-end integration latency must be reported separately from prepared-step latency.
