Runtime Troubleshooting Runbook
Collect diagnostic evidence and investigate Runtime deployment incidents.
Use this runbook for first-response Runtime incidents. The goal is to collect a bounded diagnostics bundle, identify the failing class, and decide the next action without reading source code first.
This runbook is the canonical owner of support-bundle collection order and the first decision table. Use Troubleshooting only after the runbook identifies the failing area and deeper symptom analysis is needed.
Collect The Support Bundle
For local builds, collect the bundle with the diagnostics tool:
python tools\diagnostics\collect_runtime_diagnostics.py `
--repo-root . `
--build-dir build-vs2022-x64 `
--config Release `
--json-out build-vs2022-x64\runtime-diagnostics\runtime_diagnostics.json `
--md-out build-vs2022-x64\runtime-diagnostics\runtime_diagnostics.md
When a specific pack is involved, add --pack path\to\model.pack.
Attach these files to the incident:
- diagnostics JSON and Markdown
- pack or pack provenance if it can be shared
- release bundle provenance and SBOM evidence for binary/package issues
- model create options and
tdse_model_get_execution_info(...)output - host OS, CPU/GPU/FPGA device, driver/toolkit versions, and TDSE release identifier
First Decision Table
| Symptom | Check first | Likely cause | Next action |
|---|---|---|---|
| model load fails | last_error.api_kind, last_error.status, pack error token | malformed pack, ABI mismatch, unsupported pack version | Run tdse_runtime_pack_inspect_summary(...); regenerate pack from Builder; compare release provenance. |
| unsupported pack | tdse_model_create_diagnostics_t::pack_error_code | pack version or payload contract mismatch | Rebuild pack with current SDK; keep failing pack as regression input. |
| device unavailable | tdse doctor --json-out - and tdse_model_get_execution_info(...) | provider not installed, dependency missing, authorization failure, or hardware unavailable | Repair the selected device provider. TDSE never silently substitutes another compute device. |
| TDSE GPU Add-on unavailable | hosted_addons.gpu in tdse doctor and layered CUDA diagnostics | provider load failure, driver mismatch, missing Add-on entitlement, or no GPU visible to the process | Inspect host, driver, product runtime, entitlement, and qualification separately. |
| FPGA Add-on unavailable | hosted_addons.fpga in tdse doctor and hardware evidence | provider or device image absent, authorization failure, or no compatible FPGA device | Mark blocked/unavailable; do not claim FPGA release support without real hardware evidence. |
| out of memory | memory_budget_exceeded_count, memory_peak_bytes, accelerator memory fields | pack too large, mirror/device buffers exceed budget, multi-model pressure | Reduce model count, set budget deliberately, use CPU route, or split workload. |
| concurrency violation | state.step_active, thread-safety matrix, last error API | same tdse_model_t* entered concurrently | Serialize same-handle calls; use one model handle per concurrent simulation lane. |
| performance regression | active_product_backend, route reason, hist_op_kind, memory/transfer counters | CPU_BLAS thread count changed, BLAS/CUDA provider changed, or an explicit diagnostic route was used | Compare profiler evidence and route-quality report for the same shape and host. |
| numerical mismatch | precision fields, canonical product backend name, history_compute_type | FP32 history policy, diagnostic CPU route or CUDA input route difference, pack/input mismatch | Run the internal numerical-equivalence suite in an approved test-hook build; compare pack hash and host timestep sequence. |
Symptom Procedures
Model Load Failure
- Capture create diagnostics and
tdse_runtime_pack_inspect_summary(...)output. - Confirm the pack came from the same release family as the runtime package.
- Check the release bundle provenance JSON for commit, package SHA256, and installed-files manifest.
- If the pack is malformed, keep the smallest reproducer and add it to pack fuzz or compatibility tests.
Backend Unavailable
- Run
tdse doctor --json-out -from the affected installation. - Record the selected device, Add-on state, authorization result, and provider availability reported by the command.
- For CPU execution, treat
TDSE_STATUS_UNSUPPORTEDas an OpenBLAS provider failure. Repair the provider rather than retrying through another CPU path. - For GPU or FPGA execution, repair the selected Add-on or hardware. TDSE does not silently fall back to CPU.
CUDA capability diagnosis
Do not use a CPU-only build or a skipped CUDA test as evidence that the host has no CUDA GPU. Start with the build-independent host inspector:
python3 tools/validation/runtime/check_cuda_hardware_evidence.py inspect-host
python3 tools/validation/runtime/check_cuda_hardware_evidence.py inspect-host \
--build-dir build/cuda-release
The result deliberately reports five separate states: host hardware visibility, CUDA toolkit availability, TDSE build enablement, TDSE product-runtime execution, and release qualification. Only the first state answers whether a GPU is visible; the other four must never be substituted for it.
- Confirm
hosted_addons.gpu.stateisreadyintdse doctor. - Run the CUDA hardware evidence check on the actual GPU host.
- Record driver, CUDA toolkit, GPU model, and device-provider evidence JSON.
- If evidence is missing, classify CUDA support as not certified for this release commit; do not reclassify host hardware as absent.
FPGA Unavailable
- Confirm
hosted_addons.fpga.stateintdse doctor. - Confirm the intended FPGA provider package and device image are installed.
- Run the provider readiness or hardware evidence check on the target device.
- If hardware is missing, record blocked evidence; do not substitute mock evidence for real hardware evidence.
Out Of Memory
- Inspect
memory_current_bytes,memory_peak_bytes,memory_budget_exceeded_count, and owner/residency fields. - Separate host representation bytes from CUDA pinned/device and FPGA device bytes.
- For CPU BLAS, inspect
cpu_blas_last_linear_mirror_bytesandcpu_blas_last_linear_decision. - Reduce concurrent model count, pin a lower-memory backend, or raise the budget only after capacity planning.
Performance Regression
- Capture selected device, precision, CPU thread count when applicable, and accelerator transfer counters.
- Compare against the target-host profiler report for the same
np,nh, precision, selected device, and CPU thread count when applicable. - If CPU execution is slower than expected, use the profiler on the same host before changing the host-owned thread assignment.
- If GPU transfers grew, verify that the application is still using its intended CPU-resident or GPU-resident integration mode.
Numerical Mismatch
- Reproduce with FP64 history precision.
- Reproduce with the internal numerical-equivalence suite in a test-hook build.
- Compare pack hash, timestep sequence, accepted-step count, and input/output buffer layout.
- If only FP32 differs, treat it as an application accuracy decision rather than a backend availability issue.
Escalation Criteria
Escalate to source-level debugging only after the support bundle identifies one of these:
- stable pack and same TDSE release, but reproducible runtime status failure
tdse doctorreports the selected device ready, but model creation fails- memory budget fields contradict observed allocation behavior
- release provenance, SBOM, or installed-files manifest does not match the binary package under test
- numerical mismatch reproduces on FP64 CPU reference-style execution
