Time-Domain System Equivalent logoTime-Domain System EquivalentLinear dynamics, solved faster.Discuss an evaluation
SDK Documentation

Threading and Memory Scaling

Threading, memory scaling, and operational sizing guidance.

Use this section when the one-handle rule is already settled and the next question is scale: how much memory will N models use, when does NUMA matter, when is GPU sharing still reasonable, and which model sizes deserve extra caution.

Related Chapters For concurrency rules governing individual handles, see Concurrency and Shutdown. For multi-model deployment patterns, see Multi-Model Patterns.

EDA integrations often run many TDSE models in parallel to maximize throughput. This chapter is about capacity planning and resource envelopes, not about same-handle ownership policy.

Capacity Planning For N Parallel Models

Per-Model Memory Footprint

Each TDSE model handle has a memory footprint that depends on:

  • Input-port count (np) and output-equation count (nq)
  • History depth (nh) and the packed H representation
  • Requested precision, residency, and device
  • Runtime-owned resident and preparation scratch buffers

Think about the footprint in layers:

  • Model metadata: small fixed overhead; its exact size is an implementation detail, so do not budget this category from a universal number
  • History operator storage: scales with nh * nq * np for an uncompressed dense representation
  • History and input storage: scales with the selected representation and primary/output dimensions
  • Runtime scratch buffers: preparation- and device-route-dependent

Linear Scaling Model

For N parallel model handles, memory scales approximately linearly:

Total memory approximately N x (per-model footprint) + shared_runtime_overhead

The shared process overhead is implementation- and provider-dependent. Do not budget it from a universal number.

Measure The Actual Model

Call tdse_model_get_memory_preflight(...) before execution preparation (or the first step). It reports Runtime-owned current, target, and additional preparation bytes for the selected route, and explicitly identifies estimates that are bounds or lower bounds. It does not claim allocator, driver, or other provider-private memory.

Ways To Reduce The Footprint

To reduce memory footprint with many parallel models:

  1. Reuse one model sequentially when throughput is not the goal: for repeated sweeps of the same pack, one model plus tdse_model_reset(...) can be cheaper than many parallel copies. See Multi-Model Patterns.
  2. Limit history depth deliberately: nh increases the packed history operator and history-work footprint.
  3. Choose precision and residency deliberately: use the selected model route's preflight result instead of assuming a CPU and accelerator use the same storage.

What To Measure First

Use the following to monitor memory:

  • tdse_model_info(...) returns model metadata including dimensions
  • tdse_model_get_memory_preflight(...) reports Runtime-owned capacity data
  • OS-level tools (Task Manager, top, ps) for process memory
  • For GPU workloads, use CUDA tools (nvidia-smi) to monitor GPU memory

NUMA-Aware Allocation Strategy

Current NUMA Support

The default model options leave memory placement to the host and OS. Linux also supports explicit, opt-in row-local placement for Runtime-owned CPU matrix storage through options.cpu_memory_policy = TDSE_CPU_MEMORY_LOCAL. The host first supplies the exact worker CPU order in options.cpu_ids and sets options.cpu_threads to the same count. During execution preparation, Runtime reproduces OpenBLAS's output-row task boundaries, assigns each contiguous row group to that task's NUMA node, binds only the corresponding H pages, fills them, and verifies sampled page locations. An unsupported or unverified explicit request fails before timing.

This option does not let Runtime change provider threads or CPU affinity, and it does not change the process-wide NUMA policy. Allocation, binding, page faulting, and verification all finish before the step hot path. Windows keeps the host-default memory policy.

Best Practices For NUMA Systems

For optimal performance on multi-node NUMA systems:

  1. Keep rows contiguous per node: do not round-robin adjacent H rows across dies or NUMA nodes.
  2. Match worker order and row order: options.cpu_ids are in provider worker-index order, including the caller as the last worker index.
  3. Keep host ownership explicit: the host establishes thread count and worker affinity before model preparation.
  4. Keep complete rows local: the benchmark balances contiguous row tasks across NUMA x LLC domains and binds Linux row-group pages to their task NUMA node.
// On Linux with numactl
// Run on NUMA node 0 with memory from node 0
numactl --cpunodebind=0 --membind=0 ./your_simulation

// Or programmatically:
numa_set_preferred(0);  // Prefer node 0
// Create model handles here

When NUMA Usually Matters

  • NUMA effects become visible when H no longer fits comfortably in one locality domain's cache or one node's memory bandwidth.
  • Row-local placement localizes the dominant H reads. The shared history/input vector and at most one page at each row-group boundary can still be read remotely.
  • Record the topology-derived task plan and compare target-machine matrix results; do not infer locality from thread count alone.

GPU Multi-Stream Parallelism

Current GPU Parallelism Model

Think of GPU sharing in two layers:

  • Model ownership: a Runtime model has one owner at a time, even when multiple models share a GPU.
  • Queue ownership: device-resident integration can borrow one caller-owned queue through the create options. CPU-resident integration does not expose a caller-owned CUDA queue.
  • Scheduling: TDSE does not provide a host-facing multi-stream scheduler; the host coordinates concurrently scheduled model handles and their queues.

Multi-Model GPU Parallelism

For multiple model handles on GPU:

  • Multiple model handles can be scheduled on the same selected GPU subject to their explicit resource configuration and available memory.
  • GPU memory capacity is shared across the models and other process users of that device.

GPU Memory Management

GPU memory considerations:

  • Per-model GPU memory: Matrix + solver workspace + convolution buffers
  • Shared GPU memory: CUDA context overhead, driver overhead
  • Memory limits: Use nvidia-smi to monitor GPU memory usage
  • Out-of-memory handling: TDSE returns TDSE_STATUS_OUT_OF_MEMORY if GPU allocation fails

Recommendations

  1. Monitor GPU memory: Use nvidia-smi to ensure sufficient GPU memory for your model count
  2. Reuse compatible models: When one configured Runtime model can serve a sequential workload, reuse it instead of creating concurrent duplicate handles. Circuit AC-sweep scheduling is outside the Runtime API.
  3. Prefer CPU for very small models: Small models may not benefit from GPU overhead
  4. Limit concurrent GPU models: If GPU memory is constrained, limit the number of concurrent GPU models

Practical Limits On Port Count And History Depth

Port Count Limits

Runtime represents a dense history operator with dimensions nh x nq x np. There is no customer-facing port-count recommendation that applies to every pack and device. Plan capacity from the actual np, nq, nh, precision, residency, and preflight result, then measure the target route.

History Depth Limits

For all models with an H tensor, history depth affects:

  • Memory: Linear scaling with history depth and both output/input dimensions
  • Performance: Convolution cost scales with history depth
  • Representation fidelity: The history must be long enough for the response that was qualified upstream; Runtime does not prescribe a universal depth or an accuracy threshold.

Choose the depth from the qualified model, then compare the intended target workload with a reference run. The relevant cost and capacity inputs are reported by the model's memory preflight rather than by fixed sample bands.

Solver Backend Interaction

Port count and history depth interact with solver backend selection:

  • Runtime history path: nh, nq, np, precision, and residency determine the Runtime-owned footprint and history work.
  • GPU routes: benchmark the target model and host queue behavior; TDSE does not publish a universal crossover size.

When To Expect Issues

Watch for these warning signs:

  • Memory spikes: Sudden large memory increases with small parameter changes
  • Performance degradation: Non-linear performance drop with increasing size
  • GPU out-of-memory: CUDA allocation failures

Diagnostics to monitor:

  • tdse_model_get_execution_info(...): selected and observed resources
  • tdse_model_get_memory_preflight(...): Runtime-owned capacity data before preparation
  • host and provider diagnostics appropriate to the selected deployment

Summary Checklist

Use this checklist when you are sizing a deployment, not when you are debugging a same-handle race:

  1. Estimate per-model memory: Use port count, history depth, and backend selection
  2. Plan for N x scaling: Total memory approximately N x per-model footprint
  3. Consider NUMA: On multi-socket systems, bind threads to NUMA nodes
  4. Monitor GPU memory: Use nvidia-smi for GPU workloads
  5. Start with conservative limits: Begin with moderate port counts and history depths
  6. Use diagnostics: Monitor solver_backend, matrix metrics, and numerical health
  7. Profile scaling: Test with your actual workload to verify scaling behavior