Time-Domain System Equivalent logoTime-Domain System EquivalentLinear dynamics, solved faster.Discuss an evaluation
SDK Documentation

Multi-Model Deployment Patterns

Patterns for deploying multiple TDSE models inside one host application.

Use this section when one model is no longer enough and you need to choose a deployment pattern: many independent models, one reused model for repeated sweeps, parallel copies for throughput, or separate models running at different rates.

Related Chapters For one-handle ownership and shutdown rules, see Concurrency and Shutdown. For threading and memory scaling details, see Threading and Scaling. For variable time-stepping in multi-rate scenarios, see Variable Time-Step Integration.

Use these patterns when a deployment runs different TDSE models for different subsystems, or multiple copies of the same model for parameter sweeps.

Choose A Deployment Pattern

SituationRecommended PatternWhy
independent subsystems that can run side by sideN parallel independent modelssimplest ownership story and highest operational clarity
same pack, repeated sweeps, memory is tightbatch sweep with one reused modellowest memory footprint
same pack, repeated sweeps, throughput matters mostbatch sweep with parallel copieseasiest way to scale wall-clock throughput
subsystems need different dt valuesmulti-rate couplingkeeps each model at the rate it actually needs

N Parallel Independent Models

This is the default production pattern. Each model has its own handle, its own state, and a clear owner.

#define N_MODELS 8
tdse_model_t* models[N_MODELS];

for (int i = 0; i < N_MODELS; i++) {
    tdse_model_create_diagnostics_t diag = tdse_model_create_diagnostics_init();
    tdse_model_create(
        packs[i], pack_sizes[i], NULL, &diag, &models[i]);
}

/* Each worker thread owns one model */
#pragma omp parallel for
for (int i = 0; i < N_MODELS; i++) {
    for (uint64_t n = 0; n < nsteps; ++n) {
        uint64_t hr_required = 0;
        tdse_step_begin(models[i], t[n], dt);
        tdse_step_op(models[i], &op);
        tdse_step_hr(models[i], hr, nq[i], &hr_required);
        tdse_step_commit(models[i], primary[i], np[i]);
    }
}

Rules:

  • Apply the one-handle rule from Concurrency and Shutdown: one worker owns one live handle at a time.
  • Each handle has its own history buffer and state machine.
  • Models may use different packs, different dt, devices, or precisions.

Per-Model Resource Budget

Each model handle consumes its own resources. When planning a deployment:

ResourcePer ModelShared
History and runtime storagereported per handle by memory preflightno
Operator workspaceroute- and representation-dependent; inspect memory preflightno
GPU queue/contexta caller-owned queue is required only for device-resident execution; driver/context memory is provider-privateGPU device memory and provider-private allocations are shared
CPU provider threadsone host-owned provider configurationprocess/provider scope
Execution selectionimmutable create optionsprovider availability

The provider configuration is process-owned, so concurrent CPU models must agree on the host's provider setting or run in separate processes. Record the expected provider count in each model's create options:

/* Assign CPU threads per model */
for (int i = 0; i < N_MODELS; i++) {
    tdse_model_create_options_t options = tdse_model_create_options_init();
    options.cpu_threads = provider_threads;
    tdse_model_create(
        packs[i], sizes[i], &options, NULL, &models[i]);
}

GPU Sharing Across Models

Multiple models can share the same GPU. A GPU-resident host supplies the stream for each model; device memory is shared.

Guidelines:

  • Sum the per-model GPU payloads from memory preflight, then reserve and measure additional driver and provider-private memory. The GPU preflight payload is a lower bound when those opaque allocations are present; TDSE does not publish a universal overhead number.
  • Monitor with nvidia-smi during initial deployment testing.
  • A GPU allocation failure can return TDSE_STATUS_OUT_OF_MEMORY from create, preparation, or a step, depending on when the allocation is needed.
  • Use one caller-owned stream per concurrently scheduled GPU-resident model:
for (int i = 0; i < N_MODELS; i++) {
    tdse_model_create_options_t options = tdse_model_create_options_init();
    options.device = TDSE_DEVICE_GPU;
    options.data_location = TDSE_DATA_DEVICE;
    options.device_queue = model_streams[i];
    tdse_model_create(
        packs[i], sizes[i], &options, NULL, &models[i]);
}

Batch Sweep Pattern

Use this pattern when the mathematical model stays the same but inputs, operating points, or sweep values change.

For parameter sweeps where the same pack structure is reused with different inputs:

tdse_model_t* base_model;
tdse_model_create(pack, pack_size, NULL, &diag, &base_model);

/* Option A: Sequential reuse with reset between sweeps */
for (int sweep = 0; sweep < N_SWEEPS; sweep++) {
    for (uint64_t n = 0; n < nsteps; ++n) {
        tdse_step_begin(base_model, t[n], dt);
        /* ... solve with sweep-specific primary vectors ... */
        tdse_step_commit(base_model, primary_sweep[sweep], np);
    }
    tdse_model_reset(base_model);  /* clear committed history for next sweep */
}
/* Option B: Parallel sweep with one model per sweep value */
tdse_model_t* sweep_models[N_SWEEPS];
for (int s = 0; s < N_SWEEPS; s++) {
    tdse_model_create(
        pack, pack_size, NULL, &diag, &sweep_models[s]);
}

Read the tradeoff plainly:

  • Option A is memory-efficient and simpler to operate.
  • Option B is throughput-efficient and easier to spread across workers or devices.
  • If repeated sweeps are frequent but not latency-sensitive, start with Option A.

Multi-Rate Coupling

When different subsystems require different time resolutions, use separate models with different model_dt values and synchronize at coupling boundaries:

tdse_model_t* fast;  /* model_dt = 1 ns */
tdse_model_t* slow;  /* model_dt = 10 ns */

for (step = 0; step < TOTAL_STEPS; step++) {
    tdse_step_begin(fast, t_fast, 1e-9);
    /* ... step fast model ... */
    tdse_step_commit(fast, primary_fast, fast_np);

    if (step % 10 == 0) {
        /* Extract coupling variables from fast model */
        /* ... */
        tdse_step_begin(slow, t_slow, 10e-9);
        /* ... step slow model with coupled inputs ... */
        tdse_step_commit(slow, primary_slow, slow_np);
    }

    t_fast += 1e-9;
    t_slow = (step / 10) * 10e-9;
}

See Variable Time-Step Integration for more details on multi-rate patterns.

Deployment Checklist

  • Estimate total memory: N * per-model footprint + shared overhead
  • Record one host-owned CPU-provider configuration that every concurrent CPU model can satisfy, or isolate incompatible configurations in separate processes
  • Select and benchmark the target CPU or GPU route per model; TDSE does not publish a universal size crossover
  • Verify GPU memory budget if using CUDA backend
  • Choose sweep strategy: sequential with reset vs. parallel copies
  • Test scaling: run with 1, 2, 4, 8 models and measure throughput per model
  • Monitor guard metrics on each model independently
  • Keep shutdown ownership clear: destroy each model on its owning thread or wrapper path