Time-Domain System Equivalent logoTime-Domain System EquivalentLinear dynamics, solved faster.Discuss an evaluation
SDK Documentation

Runtime Troubleshooting Runbook

Collect diagnostic evidence and investigate Runtime deployment incidents.

Use this runbook for first-response Runtime incidents. The goal is to collect a bounded diagnostics bundle, identify the failing class, and decide the next action without reading source code first.

This runbook is the canonical owner of support-bundle collection order and the first decision table. Use Troubleshooting only after the runbook identifies the failing area and deeper symptom analysis is needed.

Collect The Support Bundle

For local builds, collect the bundle with the diagnostics tool:

python tools\diagnostics\collect_runtime_diagnostics.py `
  --repo-root . `
  --build-dir build-vs2022-x64 `
  --config Release `
  --json-out build-vs2022-x64\runtime-diagnostics\runtime_diagnostics.json `
  --md-out build-vs2022-x64\runtime-diagnostics\runtime_diagnostics.md

When a specific pack is involved, add --pack path\to\model.pack.

Attach these files to the incident:

  • diagnostics JSON and Markdown
  • pack or pack provenance if it can be shared
  • release bundle provenance and SBOM evidence for binary/package issues
  • model create options and tdse_model_get_execution_info(...) output
  • host OS, CPU/GPU/FPGA device, driver/toolkit versions, and TDSE release identifier

First Decision Table

SymptomCheck firstLikely causeNext action
model load failslast_error.api_kind, last_error.status, pack error tokenmalformed pack, ABI mismatch, unsupported pack versionRun tdse_runtime_pack_inspect_summary(...); regenerate pack from Builder; compare release provenance.
unsupported packtdse_model_create_diagnostics_t::pack_error_codepack version or payload contract mismatchRebuild pack with current SDK; keep failing pack as regression input.
device unavailabletdse doctor --json-out - and tdse_model_get_execution_info(...)provider not installed, dependency missing, authorization failure, or hardware unavailableRepair the selected device provider. TDSE never silently substitutes another compute device.
TDSE GPU Add-on unavailablehosted_addons.gpu in tdse doctor and layered CUDA diagnosticsprovider load failure, driver mismatch, missing Add-on entitlement, or no GPU visible to the processInspect host, driver, product runtime, entitlement, and qualification separately.
FPGA Add-on unavailablehosted_addons.fpga in tdse doctor and hardware evidenceprovider or device image absent, authorization failure, or no compatible FPGA deviceMark blocked/unavailable; do not claim FPGA release support without real hardware evidence.
out of memorymemory_budget_exceeded_count, memory_peak_bytes, accelerator memory fieldspack too large, mirror/device buffers exceed budget, multi-model pressureReduce model count, set budget deliberately, use CPU route, or split workload.
concurrency violationstate.step_active, thread-safety matrix, last error APIsame tdse_model_t* entered concurrentlySerialize same-handle calls; use one model handle per concurrent simulation lane.
performance regressionactive_product_backend, route reason, hist_op_kind, memory/transfer countersCPU_BLAS thread count changed, BLAS/CUDA provider changed, or an explicit diagnostic route was usedCompare profiler evidence and route-quality report for the same shape and host.
numerical mismatchprecision fields, canonical product backend name, history_compute_typeFP32 history policy, diagnostic CPU route or CUDA input route difference, pack/input mismatchRun the internal numerical-equivalence suite in an approved test-hook build; compare pack hash and host timestep sequence.

Symptom Procedures

Model Load Failure

  1. Capture create diagnostics and tdse_runtime_pack_inspect_summary(...) output.
  2. Confirm the pack came from the same release family as the runtime package.
  3. Check the release bundle provenance JSON for commit, package SHA256, and installed-files manifest.
  4. If the pack is malformed, keep the smallest reproducer and add it to pack fuzz or compatibility tests.

Backend Unavailable

  1. Run tdse doctor --json-out - from the affected installation.
  2. Record the selected device, Add-on state, authorization result, and provider availability reported by the command.
  3. For CPU execution, treat TDSE_STATUS_UNSUPPORTED as an OpenBLAS provider failure. Repair the provider rather than retrying through another CPU path.
  4. For GPU or FPGA execution, repair the selected Add-on or hardware. TDSE does not silently fall back to CPU.

CUDA capability diagnosis

Do not use a CPU-only build or a skipped CUDA test as evidence that the host has no CUDA GPU. Start with the build-independent host inspector:

python3 tools/validation/runtime/check_cuda_hardware_evidence.py inspect-host
python3 tools/validation/runtime/check_cuda_hardware_evidence.py inspect-host \
  --build-dir build/cuda-release

The result deliberately reports five separate states: host hardware visibility, CUDA toolkit availability, TDSE build enablement, TDSE product-runtime execution, and release qualification. Only the first state answers whether a GPU is visible; the other four must never be substituted for it.

  1. Confirm hosted_addons.gpu.state is ready in tdse doctor.
  2. Run the CUDA hardware evidence check on the actual GPU host.
  3. Record driver, CUDA toolkit, GPU model, and device-provider evidence JSON.
  4. If evidence is missing, classify CUDA support as not certified for this release commit; do not reclassify host hardware as absent.

FPGA Unavailable

  1. Confirm hosted_addons.fpga.state in tdse doctor.
  2. Confirm the intended FPGA provider package and device image are installed.
  3. Run the provider readiness or hardware evidence check on the target device.
  4. If hardware is missing, record blocked evidence; do not substitute mock evidence for real hardware evidence.

Out Of Memory

  1. Inspect memory_current_bytes, memory_peak_bytes, memory_budget_exceeded_count, and owner/residency fields.
  2. Separate host representation bytes from CUDA pinned/device and FPGA device bytes.
  3. For CPU BLAS, inspect cpu_blas_last_linear_mirror_bytes and cpu_blas_last_linear_decision.
  4. Reduce concurrent model count, pin a lower-memory backend, or raise the budget only after capacity planning.

Performance Regression

  1. Capture selected device, precision, CPU thread count when applicable, and accelerator transfer counters.
  2. Compare against the target-host profiler report for the same np, nh, precision, selected device, and CPU thread count when applicable.
  3. If CPU execution is slower than expected, use the profiler on the same host before changing the host-owned thread assignment.
  4. If GPU transfers grew, verify that the application is still using its intended CPU-resident or GPU-resident integration mode.

Numerical Mismatch

  1. Reproduce with FP64 history precision.
  2. Reproduce with the internal numerical-equivalence suite in a test-hook build.
  3. Compare pack hash, timestep sequence, accepted-step count, and input/output buffer layout.
  4. If only FP32 differs, treat it as an application accuracy decision rather than a backend availability issue.

Escalation Criteria

Escalate to source-level debugging only after the support bundle identifies one of these:

  • stable pack and same TDSE release, but reproducible runtime status failure
  • tdse doctor reports the selected device ready, but model creation fails
  • memory budget fields contradict observed allocation behavior
  • release provenance, SBOM, or installed-files manifest does not match the binary package under test
  • numerical mismatch reproduces on FP64 CPU reference-style execution