Circuit Solver Performance
Circuit solver behavior, backend policy, and performance diagnostics.
This chapter explains Circuit performance behavior and diagnostics. Availability is a capability fact; authorization and qualification are separate product facts and must not be inferred from implementation discovery.
Runtime-managed solver selection
When you run a frequency sweep, TDSE selects the concrete solver for the installed product, authorized device class, circuit size, and circuit structure. Customer code supplies the circuit and requested operation; it does not route work to a concrete implementation.
Standard complete TDSE installations containing TDSE Circuit carry the
reviewed OpenBLAS/LAPACK dense provider and the pinned KLU sparse provider.
KLU is built into TDSE Circuit and uses the same TDSE-CIRCUIT entitlement; it
is not another product or SKU. Runtime-only installations contain neither
Circuit solver. The Runtime selects among providers recorded in the installed
capability identity. Concrete provider names may appear in diagnostics so
support teams can identify what executed; they are not customer configuration
choices.
The packaged SuiteSparse libraries are dynamically linked. A customer may
place an interface-compatible five-library replacement closure in an absolute
directory and set TDSE_LGPL_LIBRARY_OVERRIDE_DIR before the process loads a
Circuit simulation plugin. Such a replacement is outside TDSE's qualification
for the shipped library bytes and must be identified in support evidence.
All public request initializers leave the reserved solver-policy field null. Keep it null. Use returned diagnostics to record the selected device and implementation when investigating performance or a support issue.
Adaptive Sweep Planning
Choosing a frequency grid is one of the hardest parts of frequency-domain modeling. Too few points and you miss resonant features. Too many and you waste computation, or worse, produce an over-resolved spectrum that creates fitting artifacts in Builder.
The adaptive sweep planner automates this. You give it a circuit and a target time
step, and it runs planning sweeps to determine the required frequency range and
resolution. It evaluates impulse response convergence and recommends production
parameters: nfreq, nh, and dt.
#include <tdse/circuit/planning.h>
tdse_circuit_adaptive_sweep_config_t plan_cfg =
tdse_circuit_adaptive_sweep_config_init();
plan_cfg.tolerance_mode = TDSE_CIRCUIT_ADAPTIVE_SWEEP_TOLERANCE_MEDIUM;
plan_cfg.dt_target = 1e-4;
plan_cfg.limit_f_max_hz = 1e4;
// Create a report handle to capture detailed planning output
tdse_circuit_adaptive_sweep_report_t* report = NULL;
tdse_circuit_adaptive_sweep_report_create(&report);
tdse_circuit_adaptive_sweep_request_t plan_req =
tdse_circuit_adaptive_sweep_request_init();
plan_req.handle = result.handle;
plan_req.ports = ports;
plan_req.port_count = port_count;
plan_req.matrix_kind = TDSE_CIRCUIT_MATRIX_Y;
plan_req.correction_method = TDSE_BUILDER_CORRECTION_RECONSTRUCT_FROM_REAL;
plan_req.circuit_options = opt;
plan_req.planning_config = &plan_cfg;
plan_req.out_report = report;
tdse_circuit_adaptive_sweep_result_t plan_result =
tdse_circuit_adaptive_sweep_result_init();
int rc = tdse_circuit_plan_adaptive_sweep(&plan_req, &plan_result);
// plan_result.summary.production_dt_s -recommended time step
// plan_result.summary.production_nh -recommended history length
// plan_result.summary.production_nfreq -recommended frequency count
The report handle gives you detailed justification for each recommendation:
tdse_circuit_adaptive_sweep_summary_t report_summary =
tdse_circuit_adaptive_sweep_summary_init();
tdse_circuit_adaptive_sweep_report_get_summary(report, &report_summary);
// Export the full report as JSON or text for archival
size_t json_required = 0U;
tdse_circuit_adaptive_sweep_report_json(
report, NULL, 0U, &json_required);
// Allocate json_required bytes, then call again with that buffer.
size_t text_required = 0U;
tdse_circuit_adaptive_sweep_report_text(
report, NULL, 0U, &text_required);
// Allocate text_required bytes, then call again with that buffer.
tdse_circuit_adaptive_sweep_report_destroy(report);
The tolerance mode (LOOSE, MEDIUM, STRICT) controls how aggressively the
planner refines the grid. MEDIUM is the right starting point. Use STRICT when
accuracy matters more than planning time.
Tail vs. Nfreq Analysis
When preparing for Builder handoff, you may need to select an adequate frequency
resolution for a fixed frequency band. A finer grid reduces delta_f and lengthens
the reconstructed time window; that makes tail behavior observable without changing
the requested upper frequency.
The tail-vs-nfreq analysis scans both Y and Z impulse reconstructions as nfreq
increases. It reports threshold and causality metrics and recommends nfreq and
delta_f for the preferred representation. Read nh and dt from each recorded
sample, or use the adaptive planner when a complete production parameter set is
needed.
tdse_circuit_tail_convergence_request_t tail_req =
tdse_circuit_tail_convergence_request_init();
tdse_circuit_tail_convergence_config_t tail_cfg =
tdse_circuit_tail_convergence_config_init();
tail_cfg.f_max_hz = 1e4;
tail_cfg.target_nfreq = 257U;
tail_req.handle = result.handle;
tail_req.ports = ports;
tail_req.port_count = port_count;
tail_req.circuit_options = opt;
tail_req.scan_config = &tail_cfg; // required fixed-band scan configuration
// The scan evaluates both Y and Z and uses RECONSTRUCT_FROM_REAL internally.
tdse_circuit_tail_convergence_result_t tail_result =
tdse_circuit_tail_convergence_result_init();
int rc = tdse_circuit_analyze_tail_convergence(&tail_req, &tail_result);
// tail_result.summary.recommended_nfreq_preferred -recommended grid size
// tail_result.summary.recommended_delta_f_hz_preferred -recommended spacing
// tail_result.summary.preferred_threshold_achieved -threshold outcome
As with adaptive sweep planning, you can create a report handle
(tdse_circuit_tail_convergence_report_t) for per-point detail and JSON/text
export.
NPORT
NPORT is a SPICE extension that imports frequency-dependent multiport Y- or Z-parameter data from Touchstone files directly into an AC circuit solve:
NPORT NP1 1 0 2 0 FILE=example.y2p TYPE=Y
When the circuit compiles a netlist containing NPORT elements, it reads and validates the referenced Touchstone file. For each AC frequency point, the solver interpolates the data onto that point and stamps the frequency-dependent admittance block into the MNA matrix. This is automatic; no separate import or stamping API is needed.
NPORT accepts Y and Z datasets, including .ynp, .znp, .y2p, .z2p, and
their multi-port variants. An S dataset is not accepted by NPORT; use the
dedicated S-parameter element path when appropriate. The file must be readable
relative to the netlist base directory at compile time.
NPORT is an AC/frequency-domain element. Use it with matrix and AC probes;
transient series and transient probes do not provide an NPORT time-domain
model. For a runnable example, see nport_y2p_ac in the
Examples Guide.
Deep Reference: Circuit Lookup
Use this section as a lookup area after the main task sections above. It is not the best place to start if you are still deciding which circuit operation you need.
Parameter Cookbook
This table is a quick lookup for every CLI flag and its purpose. Bookmark it.
| Parameter | Where | Notes |
|---|---|---|
|text|stdin | netlist-consuming commands | file for scripts, stdin for pipes, text for inline |
--matrix y|z | matrix | Y = admittance (Norton form), Z = impedance (Thevenin form) |
--ports "p,n;p,n" | matrix, series | Order matters - it must match Builder input order |
--w0, --dw, --nfreq | matrix | Radian-frequency grid; archive it with your output |
--dc-policy | matrix | Affects DC stability: see the matrix workflow |
--series voc|isc | series | voc = open-circuit voltage, isc = short-circuit current |
--method transient|ifft|tone | series, transient probe | transient = general; ifft = freq; tone = steady-state |
--dt, --steps | series, transient probe | Must match Builder's dt for consistent handoff |
--nfft | series, probe | FFT size. ifft requires an explicit even value of at least 2; choose it for the required synthesis resolution and time window. |
--observe "v(1);i(r1)" | probe | Semicolon-separated probe expressions |
--observe-source | probe | argument = CLI only; netlist = .probe directives; argument_or_study = both. Default auto selects both when --observe is present, otherwise netlist only. |
--sweep-kind | AC probe | from_study, lin, dec, oct, list |
--ac-excitation | AC probe | small_signal for normalized transfer functions; operating_point for actual steady-state voltages |
--json-out - | circuit router | Machine-readable envelope; send data outputs to files to avoid a stdout conflict |
Terminology Map
The same concept may appear under different names across the C API, CLI, JSON/Builder outputs, and narrative documentation. Use this table when moving between surfaces:
| Concept | C API | CLI | Builder / JSON | Docs |
|---|---|---|---|---|
| Frequency sweep | compute_port_fsweep | matrix | - | "frequency sweep (fsweep)" |
| Port count | port_count | --ports (counted from list) | np | "number of ports" |
| Row count (matrix) | row_count | - | nq | "matrix rows" |
| Column count (matrix) | col_count | - | np | "matrix columns" |
| Start frequency (rad/s) | grid.w0 | --w0 | w0 | "start frequency" |
| Frequency step (rad/s) | grid.dw | --dw | dw | "frequency step / delta omega" |
| Number of frequencies | grid.nfreq | --nfreq | nfreq | "frequency count" |
| Matrix kind | TDSE_CIRCUIT_MATRIX_Y / TDSE_CIRCUIT_MATRIX_Z | --matrix y|z | - | "Y (admittance)" / "Z (impedance)" |
| Port response kind | TDSE_CIRCUIT_PORT_RESPONSE_VOC / TDSE_CIRCUIT_PORT_RESPONSE_ISC | --series voc|isc | - | "open-circuit voltage / short-circuit current" |
| Solve options | tdse_circuit_options_t | --dc-policy | - | "options / DC policy" |
| Frequency grid helper | grid_from_hz() | - | - | "Hz to rad/s conversion" |
| Probe domain | TDSE_CIRCUIT_PROBE_DOMAIN_AC / TDSE_CIRCUIT_PROBE_DOMAIN_TRANSIENT | --domain ac|transient | - | "AC / transient domain" |
Naming Conventions
TDSE Circuit uses a consistent suffix convention across all public structs:
*_count- logical object count (e.g.port_count,probe_count)*_len- buffer capacity in elements, not bytes (e.g.out_values_len). Forchar*output buffers,*_lencounts characters including the null terminator. Note:required_*_countfields on result structs are sizing query outputs ("how many elements needed") rather than buffer capacities
For APIs that document the two-call sizing contract, the result's
required_*_count field answers the sizing query ("how many elements do you
need?"). Do not infer that every result type populates every required-count
field after an arbitrary failure; follow the contract of the specific API. See
the two-output pattern.
Enum Reference
| Enum type | Values |
|---|---|
tdse_circuit_matrix_kind_t | TDSE_CIRCUIT_MATRIX_Y, TDSE_CIRCUIT_MATRIX_Z |
tdse_circuit_port_response_kind_t | TDSE_CIRCUIT_PORT_RESPONSE_VOC, TDSE_CIRCUIT_PORT_RESPONSE_ISC |
tdse_circuit_probe_domain_t | TDSE_CIRCUIT_PROBE_DOMAIN_AC, TDSE_CIRCUIT_PROBE_DOMAIN_TRANSIENT |
tdse_circuit_netlist_source_t | TDSE_CIRCUIT_NETLIST_SOURCE_TEXT, TDSE_CIRCUIT_NETLIST_SOURCE_FILE |
tdse_circuit_raw_source_t | TDSE_CIRCUIT_RAW_SOURCE_TEXT, TDSE_CIRCUIT_RAW_SOURCE_FILE |
tdse_circuit_raw_output_kind_t | TDSE_CIRCUIT_RAW_OUTPUT_BUFFER, TDSE_CIRCUIT_RAW_OUTPUT_FILE |
tdse_circuit_raw_unit_mode_t | TDSE_CIRCUIT_RAW_UNIT_PU, TDSE_CIRCUIT_RAW_UNIT_SI |
tdse_circuit_sweep_parallel_mode_t | TDSE_CIRCUIT_SWEEP_PARALLEL_OFF, TDSE_CIRCUIT_SWEEP_PARALLEL_AUTO, TDSE_CIRCUIT_SWEEP_PARALLEL_FORCE |
tdse_builder_correction_method_t | TDSE_BUILDER_CORRECTION_NONE, TDSE_BUILDER_CORRECTION_RECONSTRUCT_FROM_REAL|_IMAG|_MAG|_PHASE |
Failure Modes
When something goes wrong, start here:
| Symptom | Likely cause | Fix |
|---|---|---|
| Parse failure | Malformed netlist, missing include | Run with --json-out - on smallest deck |
| Ports resolve incorrectly | Wrong node names or polarity | Reduce to one port and verify the node pair |
| Singular or ill-conditioned solve | Floating nodes, invalid topology | Try extrapolate_from_positive DC policy or single-freq Y |
| VOC/ISC differs from expectation | Method mismatch, insufficient nfft | Compare transient vs ifft |
| NPORT import fails | Type mismatch, bad path, non-monotonic data | Check TYPE, dimensions, CWD for relative paths |
| RAW conversion fails "invalid argument" | Missing struct_size fields | Set all struct_size; check output_kind matches method |
| All backends show "no" in caps | Build configuration issue | Rebuild with required dependencies (KLU, CUDA, etc.) |
| "singular system while solving transient MNA" | Floating node, inductor cut-set, zero-impedance loop | Add GPAR minimum conductance; check topology for isolated nodes |
| "nonlinear transient solve failed to converge" | Newton iteration exhausted, stiff nonlinearity | Reduce timestep; tighten newton_residual_tol; increase newton_max_iterations |
| "singular system ... with switches" | Ideal switch creates instantaneous topology change | Add snubber conductance via GPAR; reduce timestep near switch events |
| Excessive rejected steps | Timestep too large for nonlinear dynamics | Reduce dt; check newton_abs_tol/newton_rel_tol are appropriate for variable scale |
Newton Tolerance Tuning
The tdse_circuit_time_options_t struct exposes four Newton solver fields.
Default values work for most circuits. Adjust them only when you see convergence
failures or excessive rejected steps.
| Field | Default | When to adjust |
|---|---|---|
newton_abs_tol | 1e-8 | Increase (1e-6) if voltages/currents span many orders of magnitude |
newton_rel_tol | 1e-6 | Increase (1e-4) for circuits with small signal swing; decrease (1e-8) for high precision |
newton_residual_tol | 1e-6 | Relax (1e-4) if convergence is slow but solution is adequate |
newton_max_iterations | 32 | Increase (64-128) for strongly nonlinear circuits (diodes, BJTs at turn-on) |
Progress Event Fields
tdse_circuit_progress_event_t is emitted during transient solves.
Key fields:
| Field | Meaning |
|---|---|
accepted_steps | Number of time steps the solver accepted and advanced |
rejected_steps | Number of steps rejected due to convergence failure or error estimate; a high ratio (>10%) signals dt tuning needed |
total_nonlinear_iterations | Cumulative Newton iterations across all steps; useful for comparing solver efficiency between configurations |
sim_time | Current simulation time reached |
dt | Current timestep in use |
Troubleshooting
When a command fails, collect these four things before asking for help:
- The exact command line you ran
- The full
--json-out -output - The smallest netlist that reproduces the problem
tdse circuit capsoutput from the same machine
To narrow down the problem yourself: verify the netlist parses ->verify one port ->verify one frequency ->add complexity. Almost all circuit bugs become obvious when you reduce to the simplest failing case. For Runtime-side issues after pack creation, continue in Troubleshooting.
Validation Checklist
Use this before handing off results to Builder or committing output to a project:
-
tdse circuit capsshows the expected Circuit capability - Minimal
matrixreplay passes on a known-good deck - Output dimensions match expectations (ports x frequency points)
- Builder handoff produces a valid
.packfile - Port order is consistent throughout the pipeline (circuit ->Builder ->Runtime)
Performance Evidence
Do not use a generic timing table as a capacity promise. Circuit performance depends on MNA dimension and density, right-hand-side count, frequency count, selected implementation, device, driver, and build configuration. The SDK's native dense and sparse solver measurement executables provide raw JSON/CSV input; the TDSE release-qualification workflow owns exploratory campaigns and release decisions.
Record the effective implementation from result.diagnostics; do not copy
internal thresholds or routing rules into host logic.
General storage scaling remains useful for planning:
- dense matrices require quadratic storage in the MNA dimension
- sparse matrices scale with nonzero count plus factorization fill-in
- frequency points are independent mathematically, but parallel efficiency depends on per-point work and worker/GPU overhead
Advanced Tuning And Performance
This section is for tuning and scale work after the basic circuit flow is already correct. It covers the key knobs that affect TDSE Circuit throughput and when it is worth touching them.
Understanding solver diagnostics
TDSE selection uses matrix shape, density, inverse or multi-RHS requirements,
provider availability, authorization, and qualified Runtime rules. Inspect
diagnostics.solver.backend and diagnostics.policy to learn what a specific
call actually used. These fields are evidence, not routing controls.
Parallel Frequency Sweeps
Frequency sweeps are mathematically independent across frequency points. TDSE chooses whether and how to parallelize them using the available CPU and device resources.
Parallel sweeps are most effective when each point has enough work to amortize
worker scheduling. Measure the actual deck and effective backend rather than
assuming a universal nfreq or matrix-size crossover.
When NOT to use parallel sweeps:
- Debugging: serial execution gives deterministic ordering
- Small or cheap per-point systems where worker overhead dominates
- Already running many independent circuit calls in parallel (nesting overhead)
Dense CUDA Batched Fast Path
The dense CUDA sweep implementation can batch eligible frequency-point chunks. Eligibility depends on the selected GPU implementation, affine-template support, frequency count, matrix dimension, right-hand-side count, and the runtime policy. The affine-template path currently starts at eight frequency points. Confirm use through call diagnostics and benchmark evidence; do not assume batching from CUDA availability alone.
The Runtime owns fast-path eligibility. Confirm use through call diagnostics and benchmark evidence; do not encode its internal conditions in host logic.
DC Policy Performance
The DC policy choice affects the number of DC-point factorization attempts at ω=0:
| Policy | Factorizations | Notes |
|---|---|---|
exact_dc_then_fallback | 1-2 | Tries exact first. Two factorizations only when singular |
regularized_exact_dc | 1 | Regularizes inductors with inductor_gbig |
extrapolate_from_positive | 0 | Skips a DC solve and derives the endpoint from positive-frequency points |
Sparse Solver Factor Caching
KLU and SPARSE_CUDA solvers cache the symbolic factorization pattern. When
computing multiple sweeps against the same circuit with different frequency
grids, the symbolic analysis is reused. This is automatic - no configuration
needed. The diagnostics field factor_cache_enabled and factor_cache_hit
report whether caching was used.
Affine AC Build
When eligible, the circuit builds an affine template of the MNA system and materializes subsequent frequency points from it instead of rebuilding the full system. The current minimum is eight frequency points, and the circuit must support affine stamping. Treat any speedup as workload-specific evidence.
Compile Once, Sweep Many
A compiled handle is reusable across multiple compute calls. Always compile once and reuse the handle for multiple sweeps, rather than recompiling:
// Good -compile once
tdse_modelspace_circuit_compile(&req, &result);
for (int k = 0; k < num_configs; ++k) {
freq_req.handle = result.handle;
freq_req.grid = grids[k];
tdse_circuit_compute_port_frequency_sweep(&freq_req, &freq_result);
}
// Bad -recompiles every iteration
for (int k = 0; k < num_configs; ++k) {
tdse_modelspace_circuit_compile(&req, &result);
freq_req.handle = result.handle;
tdse_circuit_compute_port_frequency_sweep(&freq_req, &freq_result);
tdse_circuit_destroy(&result.handle);
}
Compilation involves parsing, MNA construction, and topology analysis. Reuse avoids that work; the actual ratio to a solve is workload-dependent.
SPARSE_CUDA provider status
When the contracted installation includes the GPU Add-on, AUTO may select the
packaged sparse CUDA provider for eligible systems. Use tdse doctor for the
installed product and authorization view, and the execution-capability report
for provider/device detail. Provider presence, device visibility, and formal
qualification remain separate facts. Dependency acquisition, receipts, and
qualification-build switches are release-engineering procedures outside the
customer workflow.
