Threading and Memory Scaling
Threading, memory scaling, and operational sizing guidance.
Use this section when the one-handle rule is already settled and the next question is scale: how much memory will N models use, when does NUMA matter, when is GPU sharing still reasonable, and which model sizes deserve extra caution.
Related Chapters For concurrency rules governing individual handles, see Concurrency and Shutdown. For multi-model deployment patterns, see Multi-Model Patterns.
EDA integrations often run many TDSE models in parallel to maximize throughput. This chapter is about capacity planning and resource envelopes, not about same-handle ownership policy.
Capacity Planning For N Parallel Models
Per-Model Memory Footprint
Each TDSE model handle has a memory footprint that depends on:
- Input-port count (
np) and output-equation count (nq) - History depth (
nh) and the packedHrepresentation - Requested precision, residency, and device
- Runtime-owned resident and preparation scratch buffers
Think about the footprint in layers:
- Model metadata: small fixed overhead; its exact size is an implementation detail, so do not budget this category from a universal number
- History operator storage: scales with
nh * nq * npfor an uncompressed dense representation - History and input storage: scales with the selected representation and primary/output dimensions
- Runtime scratch buffers: preparation- and device-route-dependent
Linear Scaling Model
For N parallel model handles, memory scales approximately linearly:
Total memory approximately N x (per-model footprint) + shared_runtime_overhead
The shared process overhead is implementation- and provider-dependent. Do not budget it from a universal number.
Measure The Actual Model
Call tdse_model_get_memory_preflight(...) before execution preparation (or
the first step). It reports Runtime-owned current, target, and additional
preparation bytes for the selected route, and explicitly identifies estimates
that are bounds or lower bounds. It does not claim allocator, driver, or other
provider-private memory.
Ways To Reduce The Footprint
To reduce memory footprint with many parallel models:
- Reuse one model sequentially when throughput is not the goal: for repeated sweeps of the same pack, one model plus
tdse_model_reset(...)can be cheaper than many parallel copies. See Multi-Model Patterns. - Limit history depth deliberately:
nhincreases the packed history operator and history-work footprint. - Choose precision and residency deliberately: use the selected model route's preflight result instead of assuming a CPU and accelerator use the same storage.
What To Measure First
Use the following to monitor memory:
tdse_model_info(...)returns model metadata including dimensionstdse_model_get_memory_preflight(...)reports Runtime-owned capacity data- OS-level tools (Task Manager,
top,ps) for process memory - For GPU workloads, use CUDA tools (
nvidia-smi) to monitor GPU memory
NUMA-Aware Allocation Strategy
Current NUMA Support
The default model options leave memory placement to the host and OS. Linux
also supports explicit, opt-in row-local placement for Runtime-owned CPU
matrix storage through options.cpu_memory_policy = TDSE_CPU_MEMORY_LOCAL.
The host first supplies the exact worker CPU order in options.cpu_ids and
sets options.cpu_threads to the same count. During execution preparation,
Runtime reproduces OpenBLAS's output-row task boundaries, assigns each
contiguous row group to that task's NUMA node, binds only the corresponding H
pages, fills them, and verifies sampled page locations. An unsupported or
unverified explicit request fails before timing.
This option does not let Runtime change provider threads or CPU affinity, and it does not change the process-wide NUMA policy. Allocation, binding, page faulting, and verification all finish before the step hot path. Windows keeps the host-default memory policy.
Best Practices For NUMA Systems
For optimal performance on multi-node NUMA systems:
- Keep rows contiguous per node: do not round-robin adjacent H rows across dies or NUMA nodes.
- Match worker order and row order:
options.cpu_idsare in provider worker-index order, including the caller as the last worker index. - Keep host ownership explicit: the host establishes thread count and worker affinity before model preparation.
- Keep complete rows local: the benchmark balances contiguous row tasks across NUMA x LLC domains and binds Linux row-group pages to their task NUMA node.
Recommended NUMA Pattern
// On Linux with numactl
// Run on NUMA node 0 with memory from node 0
numactl --cpunodebind=0 --membind=0 ./your_simulation
// Or programmatically:
numa_set_preferred(0); // Prefer node 0
// Create model handles here
When NUMA Usually Matters
- NUMA effects become visible when H no longer fits comfortably in one locality domain's cache or one node's memory bandwidth.
- Row-local placement localizes the dominant H reads. The shared history/input vector and at most one page at each row-group boundary can still be read remotely.
- Record the topology-derived task plan and compare target-machine matrix results; do not infer locality from thread count alone.
GPU Multi-Stream Parallelism
Current GPU Parallelism Model
Think of GPU sharing in two layers:
- Model ownership: a Runtime model has one owner at a time, even when multiple models share a GPU.
- Queue ownership: device-resident integration can borrow one caller-owned queue through the create options. CPU-resident integration does not expose a caller-owned CUDA queue.
- Scheduling: TDSE does not provide a host-facing multi-stream scheduler; the host coordinates concurrently scheduled model handles and their queues.
Multi-Model GPU Parallelism
For multiple model handles on GPU:
- Multiple model handles can be scheduled on the same selected GPU subject to their explicit resource configuration and available memory.
- GPU memory capacity is shared across the models and other process users of that device.
GPU Memory Management
GPU memory considerations:
- Per-model GPU memory: Matrix + solver workspace + convolution buffers
- Shared GPU memory: CUDA context overhead, driver overhead
- Memory limits: Use
nvidia-smito monitor GPU memory usage - Out-of-memory handling: TDSE returns
TDSE_STATUS_OUT_OF_MEMORYif GPU allocation fails
Recommendations
- Monitor GPU memory: Use
nvidia-smito ensure sufficient GPU memory for your model count - Reuse compatible models: When one configured Runtime model can serve a sequential workload, reuse it instead of creating concurrent duplicate handles. Circuit AC-sweep scheduling is outside the Runtime API.
- Prefer CPU for very small models: Small models may not benefit from GPU overhead
- Limit concurrent GPU models: If GPU memory is constrained, limit the number of concurrent GPU models
Practical Limits On Port Count And History Depth
Port Count Limits
Runtime represents a dense history operator with dimensions nh x nq x np.
There is no customer-facing port-count recommendation that applies to every
pack and device. Plan capacity from the actual np, nq, nh, precision,
residency, and preflight result, then measure the target route.
History Depth Limits
For all models with an H tensor, history depth affects:
- Memory: Linear scaling with history depth and both output/input dimensions
- Performance: Convolution cost scales with history depth
- Representation fidelity: The history must be long enough for the response that was qualified upstream; Runtime does not prescribe a universal depth or an accuracy threshold.
Choose the depth from the qualified model, then compare the intended target workload with a reference run. The relevant cost and capacity inputs are reported by the model's memory preflight rather than by fixed sample bands.
Solver Backend Interaction
Port count and history depth interact with solver backend selection:
- Runtime history path:
nh,nq,np, precision, and residency determine the Runtime-owned footprint and history work. - GPU routes: benchmark the target model and host queue behavior; TDSE does not publish a universal crossover size.
When To Expect Issues
Watch for these warning signs:
- Memory spikes: Sudden large memory increases with small parameter changes
- Performance degradation: Non-linear performance drop with increasing size
- GPU out-of-memory: CUDA allocation failures
Diagnostics to monitor:
tdse_model_get_execution_info(...): selected and observed resourcestdse_model_get_memory_preflight(...): Runtime-owned capacity data before preparation- host and provider diagnostics appropriate to the selected deployment
Summary Checklist
Use this checklist when you are sizing a deployment, not when you are debugging a same-handle race:
- Estimate per-model memory: Use port count, history depth, and backend selection
- Plan for N x scaling: Total memory approximately N x per-model footprint
- Consider NUMA: On multi-socket systems, bind threads to NUMA nodes
- Monitor GPU memory: Use
nvidia-smifor GPU workloads - Start with conservative limits: Begin with moderate port counts and history depths
- Use diagnostics: Monitor
solver_backend, matrix metrics, and numerical health - Profile scaling: Test with your actual workload to verify scaling behavior
