Step Execution Model
Per-step runtime semantics for begin, evaluate, commit, and output handling.
Use this section when you are wiring TDSE Runtime into a simulator step loop. It explains what each step call means, how trial state differs from committed state, how retries behave when a trial is rejected, and what to record when the loop does not behave as expected.
Related Chapters For lifecycle and shutdown semantics, see Runtime Lifecycle. For concurrency rules, see Concurrency and Shutdown.
If the lifecycle section answers "who owns the handle," this section answers "what exact state is
the handle in between begin, query calls, commit, and dr?"
For most customer integrations, this section is the contract source for one of these three host patterns:
| Host pattern | What to focus on here |
|---|---|
| fixed-step solver loop | canonical step order and prime-step pattern |
| adaptive / trial-rejecting solver loop | rejected-trial semantics and state snapshots |
| real-time / HIL loop | committed-state discipline, dr timing, and one-handle-per-thread rules |
The Three Invariants
If you remember only three step-loop rules, remember these:
-
tdse_step_begin(...)creates a trial context at one exact(t, dt). -
These trial query calls are query-only within that trial context:
tdse_step_op(...)tdse_step_hr(...)tdse_step_ir(...)
-
tdse_step_commit(...)is the preferred call that advances committed history.
Everything else in this chapter follows from those three rules.
Canonical Step Order
Per ordinary step, the runtime flow is:
tdse_step_begin(model, t, dt)tdse_step_op(model, op_out)as neededtdse_step_hr(model, hr_out, hr_capacity, &hr_required)tdse_step_ir(model, ir_out, ir_capacity, &ir_required)when IR exists and is still in range- host-side trial solve or trial update
tdse_step_commit(model, primary_accepted, primary_count)if the trial is accepted- optional
tdse_step_dr(model, dr_out, dr_capacity, &dr_required)after commit
tdse_step_op(...) is an instantaneous operator query. Hosts may:
- fetch it once and cache it under LTI assumptions
- fetch it per step when host policy wants explicit freshness
Host/TDSE ownership split in this loop:
-
the host owns
t,dt, trial acceptance, solver state, and the acceptedprimaryvector -
TDSE owns trial-local query state, committed history, and the resulting
op/hr/ir/drsurfaces
Prepared Realtime Interval
For an explicitly selected CPU model with fixed-step, square, uniform-grid history, the host can move Runtime setup outside its deadline loop:
- create the model with its device, precision, and CPU resource configuration
- call
tdse_model_execution_prepare(model)to allocate and validate the selected execution resources before the first step - if first-use compute latency matters, optionally run representative trial
queries and finish each trial with
tdse_step_discard(...); this warms the prepared compute path without advancing committed history - run any physically required initialization or pre-event interval through the normal step loop, committing the accepted primary values supplied by the host
- fetch and cache
tdse_step_op(...)if operator conditions are unchanged - run only the required step calls in the deadline-critical interval
- call
tdse_model_execution_end(model)on the same owner thread after the last completed step - collect diagnostics and perform administrative cleanup outside the loop
Preparation is explicit and fail-closed. It may allocate representations,
validate immutable H/IR data, and bind the history and commit plans. A
successful return proves that work completed before the loop; it also
acquires the retained owner lease, drains older snapshot readers, and closes
the prepare request context. The first realtime
tdse_step_begin(...) therefore performs none of that work lazily.
Warmup, Computational Prewarm, and Physical Initialization
TDSE deliberately has no tdse_warmup(...) Runtime API. The word warmup is
often used for three different operations, and combining them would let the
Runtime invent circuit state that only the host can define.
| Operation | Purpose | Advances committed history? | Required? |
|---|---|---|---|
| execution preparation | create fixed representations, upload H, and prepare the selected provider | no | yes before a deadline-critical interval |
| computational prewarm | exercise the exact production query path so first-use effects occur before the deadline | no | recommended when first-step latency matters |
| physical initialization | establish a real pre-event or initial-condition history from host-accepted primary values | yes | only when required by the circuit model |
The corresponding order is:
- call
tdse_model_execution_prepare(...); - optionally prewarm with trial queries followed by
tdse_step_discard(...); - run any physically meaningful initialization steps and commit their real accepted primary values;
- enter the deadline-critical interval;
- call
tdse_model_execution_end(...)after the last completed step.
Do not call tdse_step_commit(...) with zeros or synthetic values merely to
warm the machine. A commit is a physical state transition. If zero prehistory
is the intended initial condition, no history prefill is required: delayed
samples before the oldest committed value contribute zero.
Recommended Application Prewarm
Use one rule on every supported device: execute 1024 non-committing production-path calls after preparation and before the deadline-critical interval. The count is the same for CPU, GPU with host-resident data, and GPU with device-resident data. Use the exact production device, precision, data location, queue, buffers, OpenBLAS thread setting, and CPU set.
This is a startup recommendation, not a numerical correctness requirement. The fixed count deliberately avoids a timer, adaptive loop, and separate CPU and GPU policies in customer code. For an asynchronous GPU queue, confirm that all 1024 calls have completed before entering production; merely enqueueing them is not sufficient.
Separate computational prewarm is unnecessary when the host already runs a physically valid initialization or pre-event interval through the same production calls, buffers, queue, precision, and resource assignment for at least 1024 steps. Those committed steps naturally exercise the compute path while establishing the required history.
Repeat the prewarm after recreating the model or changing the device, precision, data location, borrowed queue, CPU thread contract, CPU set, or provider. Consider repeating it after a long idle interval when processor power state or shared-cache eviction makes first-step latency important.
Host-Resident Prewarm Example
After successful preparation, repeat the next trial at the same (t, dt) and
discard it every time:
const size_t prewarm_calls = 1024;
for (size_t i = 0; i < prewarm_calls; ++i) {
tdse_status_t query_status = tdse_step_begin(model, first_t, dt);
if (query_status != TDSE_STATUS_OK) {
return query_status;
}
query_status = tdse_step_hr(model, hr, hr_capacity, &hr_required);
/* Clear the trial even when the query failed. */
tdse_status_t discard_status = tdse_step_discard(model);
if (query_status != TDSE_STATUS_OK) {
return query_status;
}
if (discard_status != TDSE_STATUS_OK) {
return discard_status;
}
}
Call the same query set that production will use. For example, do not add
tdse_step_op(...) merely for warmup when production fetches and caches the
operator once. If production consumes ir, include tdse_step_ir(...) while
it is in range. Keep timing, logging, telemetry, and stability checks outside
the eventual deadline loop.
Device-Resident GPU Prewarm
The lifecycle remains begin/query/discard, but the history query uses the device-resident boundary:
const size_t prewarm_calls = 1024;
for (size_t i = 0; i < prewarm_calls; ++i) {
tdse_status_t query_status = tdse_step_begin(model, first_t, dt);
if (query_status != TDSE_STATUS_OK) {
return query_status;
}
query_status = tdse_eval_history_device_typed(
model, device_hr, nq, scalar_bits);
tdse_status_t discard_status = tdse_step_discard(model);
if (query_status != TDSE_STATUS_OK) {
return query_status;
}
if (discard_status != TDSE_STATUS_OK) {
return discard_status;
}
}
host_confirm_queue_complete(caller_owned_queue);
host_confirm_queue_complete(...) represents the host's normal completion
mechanism for its borrowed queue; it is not a TDSE API. Completion must be
confirmed before prewarm is considered finished. Do not submit a fake device
primary vector through tdse_commit_primary_device_typed(...).
What Warmup Can and Cannot Guarantee
Preparation makes the GPU H matrix resident in device memory. Warmup can exercise kernels, graphs, buffers, queue ordering, CPU code pages, and current working data, and it can allow CPU/GPU clocks to leave an idle state. It does not pin the complete H matrix in CPU cache, GPU L2, or on-chip SRAM:
- CPU caches remain managed by the processor and can be evicted by the host;
- a small GPU working set may benefit from L2 persistence, but that is a hint, not permanent ownership;
- a large H matrix remains resident in VRAM or HBM and is streamed through cache during each prepared step.
Therefore, successful warmup means that first-use work has been exercised; it is not proof that every coefficient will remain in the fastest cache level. Hard real-time acceptance must still be based on measured deployment-hardware latency after the complete host initialization sequence.
While the prepared owner interval is active, TDSE retains one handle lifetime lease and execution guard across all step calls. It does not reacquire the process lifetime registry, rebuild request diagnostics, or wait for diagnostic snapshot readers on every sub-call. Snapshot diagnostics attempted from other threads fail bounded during this interval so they cannot delay the owner. FP32 output validation is fused with FP32-to-FP64 conversion. Monolithic FP64 retains its required numerical-finiteness scan because that result can overflow even when all prepared inputs are finite.
The interval does not make optional work mandatory. Do not call
tdse_step_op(...) every step when its conditions have not changed, do not call
tdse_step_ir(...) for a model without IR, and do not call
tdse_step_dr(...) when the host does not consume direct response.
The stable execution boundary is tdse_model_execution_prepare(...) followed
by tdse_model_execution_end(...). CPU, GPU, and FPGA hosts use this same
lifecycle; provider-specific preparation remains internal.
History Prefetch In Real-Time Loops
When the SDK is built with TDSE_ENABLE_HISTORY_PREFETCH=ON (the embedded
customer profile builds with it OFF), TDSE may internally precompute
delayed-history work for the next step while the host solver is busy. This
does not change the public call order or the same-handle ownership rule: the
host still calls the public step APIs from one owner thread per model handle.
History prefetch is a correctness-neutral optimization:
- cache hits can reduce the next
tdse_step_hr(...)latency - cache misses, incomplete workers, stale keys, route mismatches, and worker failures fall back to the synchronous delayed-history implementation
- best-effort fallback preserves numerical correctness, but it may be reported as deadline risk in real-time diagnostics
- the prefetch worker's optimized path is fixed-shape uniform history,
including rectangular
[nq][np]models - variable-dt and explicit-tau history continue through the existing interpolation path
CPU prefetch must be budgeted against the host solver. If the host and TDSE together cannot finish a step inside the real-time period, enabling prefetch cannot make that workload hard-real-time ready; it only moves eligible work into measured slack.
Prime-Step Pattern
The standard integration pattern primes at n = -1:
tdse_step_begin(m, t0 - dt, dt);
tdse_step_op(m, &op_square);
tdse_step_commit(m, primary_minus1, np);
Why this matters:
- it establishes committed history before the first ordinary simulation step
- it makes later
hrevaluation align with the expected discrete-time history model - it gives
dra valid first committed state immediately after prime
If you skip the prime step without redesigning initialization assumptions, the first visible issue often looks numerical even though the real bug is committed-history alignment.
Prime-step design review question:
- what exact accepted primary vector should the runtime treat as the prehistory state?
If that answer is still implicit, the integration is not really complete.
Mathematical Contract
The trial-side relation is:
y_trial = op * primary_trial + hr + ir
Where:
opis the instantaneous operatorhris the delayed-history contributioniris the independent-response contribution
The committed direct-response relation is:
dr[n] = op * primary_accepted[n]
Legality Matrix
This is the fastest table to use during code review or triage:
| API | Requires Live Handle | Requires Active Trial Step | Requires Prior Commit | Advances State |
|---|---|---|---|---|
tdse_step_begin(...) | yes | no, unless exact re-entry | no | creates or re-enters trial context |
tdse_step_op(...) | yes | yes | no | no |
tdse_step_hr(...) | yes | yes | no | no |
tdse_step_ir(...) | yes | yes | no | no |
tdse_step_commit(...) | yes | yes | no | yes |
tdse_step_dr(...) | yes | no | yes | no |
The public hr, ir, commit, and dr functions are the capacity/count-
checked step interface. For hr, ir, and dr, a NULL/zero sizing call
reports Nq through out_required; an undersized buffer returns
TDSE_STATUS_BUFFER_TOO_SMALL without a partial write. commit requires an
explicit count equal to Np and does not advance history when that validation
fails.
What tdse_model_state_info(...) Means During A Step Loop
The most useful runtime snapshot during integration is tdse_model_state_info(...).
Its high-value fields are:
| Field | Operational Meaning |
|---|---|
step_active | nonzero while a trial step is currently active |
has_committed_step | nonzero after at least one successful commit |
committed_steps | number of accepted commits |
committed_t / committed_dt | time coordinates of the latest committed step |
active_t / active_dt | active trial-step coordinates, or zero when no trial is active |
sim_time | accumulated committed simulation time |
dr_last_valid | whether tdse_step_dr(...) is currently queryable |
This snapshot reports dynamic execution state only. It is the quickest way to distinguish these three cases:
- no trial has started yet
- a trial is active but not committed
- a committed step exists and
dris valid
State Snapshot Transition Table
The manual previously named these fields, but host teams usually need the exact transition pattern. Use this table as the expected runtime fingerprint.
| Phase | step_active | has_committed_step | committed_steps | active_t / active_dt | committed_t / committed_dt | dr_last_valid |
|---|---|---|---|---|---|---|
| newly created handle | 0 | 0 | 0 | 0 / 0 | 0 / 0 | 0 |
after begin(t, dt) | 1 | unchanged | unchanged | t / dt | unchanged | unchanged |
| after trial queries only | 1 | unchanged | unchanged | t / dt | unchanged | unchanged |
after accepted commit(...) | 0 | 1 | increments by 1 | 0 / 0 | becomes the just-committed t / dt | 1 |
| after rejected trial with no commit | 1 until caller leaves the active trial lifecycle | unchanged | unchanged | still the active trial coordinates | unchanged | unchanged |
after dr(...) | 0 | 1 | unchanged | 0 / 0 | unchanged | 1 |
Two practical readings matter:
- if
step_active=1andcommitted_stepsis not moving, the runtime is still in a trial context - if
dr_last_valid=1, at least one accepted commit already exists and no new commit is required just to read the latest committed direct response
begin Semantics And Re-Entry
tdse_step_begin(...) does three user-visible things:
- validates that the handle and
(t, dt)are legal - on a non-prepared path, freezes execution-affecting runtime controls on the
first successful begin;
tdse_model_execution_prepare(...)freezes them before the prepared interval starts - records the active trial coordinates and marks
step_active=1
The important re-entry rule is strict:
-
re-entering
tdse_step_begin(...)while a trial is already active is valid only whentanddtmatch the already-active trial step within runtime tolerance -
mismatched re-entry fails with
TDSE_STATUS_INVALID_STATE
This means begin is not a generic "start over" button.
It is an idempotent re-entry only for the same trial coordinates.
What Re-Entry Is For
Exact re-entry is useful when:
- a host wrapper retries a query path but is still on the same trial step
- instrumentation or layered call sites may invoke the same "ensure step active" helper twice
It is not valid for:
- changing
tmid-trial - changing
dtmid-trial - silently converting a rejected trial into a different step without resolving the active state
Rejected Trials And Retry Semantics
Rejected trials are normal in real host integrations. The runtime contract is intentionally simple:
- start the trial with
begin - query
op,hr, and optionalir - host decides the trial solve is not acceptable
- do not call
commit - treat committed history as unchanged
The one rule that matters most is this:
- no
commitmeans no committed-state movement
Operational consequences:
committed_stepsdoes not increasecommitted_tandcommitted_dtdo not changesim_timedoes not advancedr_last_validstays tied to the previous committed step
How To Retry Cleanly
If the host rejects the trial and wants to try again:
- keep the retry anchored to the same step coordinates if the same trial is being reconsidered
- do not pretend a new committed step exists
- do not read the lack of state movement as a runtime failure; it is the intended rejection model
If the host instead wants a different (t, dt), it must follow the documented lifecycle rather
than mismatched begin re-entry.
Worked Snapshot Trace
This is the most useful support trace to memorize.
Before Prime
Expected snapshot:
step_active=0has_committed_step=0committed_steps=0dr_last_valid=0
Interpretation:
- no committed history exists yet
dr(...)is invalid at this point
After Prime begin(t0 - dt, dt)
Expected snapshot:
step_active=1active_t=t0 - dtactive_dt=dthas_committed_step=0
Interpretation:
- the runtime now has an active trial context
- committed history still does not exist until
commit
After Prime commit(primary_minus1)
Expected snapshot:
step_active=0has_committed_step=1committed_steps=1committed_t=t0 - dtcommitted_dt=dtdr_last_valid=1
Interpretation:
- history now exists
dr(...)is legal- the first ordinary step can use a properly anchored committed past
During Ordinary Step n
After begin(t_n, dt) and before commit(...):
step_active=1active_t=t_nactive_dt=dtcommitted_stepsstill equals the previous accepted countcommitted_tstill points to the previous accepted step
This is the point where support should ask:
- are we still in a legitimate trial?
- is the host expecting committed values too early?
After Accepted Commit At Step n
Expected snapshot:
step_active=0committed_stepsincrements by onecommitted_t=t_ncommitted_dt=dtsim_timeincreases bydtdr_last_valid=1
Interpretation:
- committed history has advanced
- future
hrcalls will now see the newly accepted state in their delayed-history basis
After Rejected Trial At Step n
If the host decides not to commit:
committed_stepsremains unchangedcommitted_tremains the previous committed timesim_timeremains unchangeddr_last_validstill refers to the previous committed step
This is the exact fingerprint of "trial rejected, committed history preserved."
Worked Host Pattern
The intended control split is:
- Runtime supplies
op,hr, and optionalir - host code computes or solves the trial equation
- host decides whether the trial primary vector is accepted
- Runtime advances history only when the host commits
This is why Runtime remains auditable inside larger simulator loops.
Minimal Integration Recipes
Use the smallest recipe that matches your host:
| Host type | Minimal recipe |
|---|---|
| fixed-step simulator | prime once, then begin -> op/hr/ir -> host solve -> commit -> dr |
| adaptive simulator | same loop, but host may reject a trial and skip commit |
| RT/HIL loop | fixed accepted-step policy, one handle per execution thread, dr only after commit |
For code review, the key question is not "did the host call the APIs" but "which side owns acceptance, timing, and history movement?"
C Example
tdse_step_begin(model, t, dt);
tdse_step_op(model, &op);
uint64_t hr_required = 0;
tdse_status_t hr_st = tdse_step_hr(model, hr, nq, &hr_required);
uint64_t ir_required = 0;
tdse_status_t ir_st = tdse_step_ir(model, ir, nq, &ir_required);
if (ir_st == TDSE_STATUS_OUT_OF_RANGE) {
/* Host policy decides whether this is a hard stop or planned IR horizon exit. */
}
if (hr_st == TDSE_STATUS_OK && ir_st == TDSE_STATUS_OK) {
solve_trial(op, hr, ir, primary_trial, y_trial);
}
if (trial_is_accepted(primary_trial, y_trial)) {
tdse_step_commit(model, primary_trial, np);
uint64_t dr_required = 0;
tdse_step_dr(model, dr, nq, &dr_required);
}
The key handbook reading is:
- the host owns acceptance
- the runtime owns committed-history mutation
Fixed-Step Host Skeleton
This is the simplest integration that should be proven before any adaptive or real-time embellishment:
prime_once(model, t0, dt, primary_minus1);
for (size_t n = 0; n < nsteps; ++n) {
tdse_step_begin(model, t0 + n * dt, dt);
tdse_step_op(model, &op);
uint64_t hr_required = 0, ir_required = 0, dr_required = 0;
tdse_step_hr(model, hr, nq, &hr_required);
tdse_step_ir(model, ir, nq, &ir_required); /* when IR exists */
host_solve_trial(op, hr, ir, primary_trial, y_trial);
tdse_step_commit(model, primary_trial, np);
tdse_step_dr(model, dr, nq, &dr_required);
}
If this path is not already stable, stop here before adding retries, adaptive
dt, or multi-threading.
Adaptive-Step Ownership Rule
For adaptive hosts, keep one rule explicit in design notes:
- the host may retry or reject a trial
- TDSE must not see history advancement unless the host actually commits
That single rule prevents most accidental state-drift bugs.
IR Horizon Behavior Inside The Loop
tdse_step_ir(...) is query-only, but it is still a contract boundary.
If the current step time lies beyond configured IR support, the API returns
TDSE_STATUS_OUT_OF_RANGE.
Read that status as:
- the runtime is telling you the packaged
IRsequence does not extend to this step - the runtime is not silently extrapolating
IR
What to record:
- archive
committed_steps,committed_t,dt, and modelir_nsteps - verify whether the host advanced beyond the packaged horizon by design or by mistake
Rectangular Versus Square Operator View
tdse_step_op(...) supports:
- square view:
np x np - full rectangular view:
nq x np
Recommended practice:
- choose one operator-view policy per integration
- document that policy in the host design
- do not switch views ad hoc across call sites without a clear reason
Using the wrong view usually appears later as a shape or interpretation bug, not at the moment the host design drifted.
What to Capture for Step Incidents
When a step-loop issue shows up, collect these details in this order:
- failing API name and returned status
tdse_model_state_info(...)tdse_model_last_error_info(...)- exact
t,dt, and host step index - whether the host intended to accept or reject the trial
- whether prime was performed
- whether the host expected square or rectangular
op
This sequence usually resolves three common ambiguities immediately:
- lifecycle misuse versus numerical rejection
- missing prime versus wrong-history interpretation
IRhorizon exit versus unrelated step failure
Common Step-Loop Mistakes
| Mistake | Why It Is Wrong | Correct Pattern |
|---|---|---|
calling dr during an active trial step | dr is post-commit only | query dr after commit, with no active trial step |
treating hr as state-advancing | hr is query-only | call commit to advance history |
| skipping prime without redesigning history assumptions | later history alignment shifts | prime at n = -1 or document a different initialization contract |
re-entering begin with different t or dt | active trial context must remain one exact step | finish the active trial lifecycle before changing coordinates |
| assuming rejected trial implies hidden state rollback logic | no commit means no advancement happened | read snapshots and keep retry logic explicit |
| mixing square and rectangular operator views ad hoc | host matrix assumptions drift | choose and document one operator-view policy |
Anti-Patterns
Avoid these integration patterns even if they seem harmless in local testing:
- using
state_infoonly after failures instead of also learning the normal expected snapshot - inferring committed advancement from successful
hrorirqueries - treating
dr_last_valid=1as proof that the current trial was committed - retrying a rejected trial by changing
(t, dt)under an already-active trial step - explaining a history bug as "numerical noise" before verifying prime and commit sequencing
