backend#

Measures C++ backend test cases whose names match an ECMAScript regular expression. test mode generates the ordinary correctness cases, while benchmark mode generates enlarged inputs for operators that provide a benchmark case:

python -m onnx_light backend \
    --regex "^test_cc_(abs|gemm)" \
    --mode benchmark

The report separates lazy case materialization, evaluator setup, warm-up, and measured execution time for every selected case. --repeat controls measured iterations and --warmup controls unmeasured iterations. Defaults are ten measured iterations and two warm-up iterations per logical CPU. Each phase is limited to one cumulative second by default; --max-repeat-time SECONDS changes this limit. Each case runs in an isolated process for at most two seconds by default. --timeout SECONDS changes that limit; timed-out cases are stopped, reported with status=timeout, and followed by the remaining cases. By default, big cases containing _big_ are excluded; --include-big includes them. --json returns the complete machine-readable report. --output writes one row per selected case as CSV or XLSX according to the file extension:

python -m onnx_light backend --regex ".*not.*" --output not.csv
python -m onnx_light backend --regex ".*not.*" --output not.xlsx

XLSX output contains a summary sheet with the run configuration and CPU descriptor, plus a backend sheet with one row per case.

--save-models DIRECTORY saves every completed test as DIRECTORY/<test-name>.onnx. The test name is also stored as the ONNX graph name. Each model is self-contained in one file; timed-out tests do not produce a model:

python -m onnx_light backend --regex ".*not.*" \
    --save-models backend-models

Use the same --parameter syntax as kernel --tune to run backend cases side by side with explicit tuning values. --kernel, --dtype and --impl select the exact tuning schema, and --criterion selects the metric to optimize:

python -m onnx_light backend --regex ".*not.*" \
    --kernel Not --dtype BOOL --impl portable \
    --parameter parallel.minimum_elements=default,16384,32768 \
    --criterion median-speedup

default resolves to the active value and defines the speedup baseline. Repeat --parameter to evaluate the Cartesian product of several parameters. Every selected backend case runs once per parameter set. The report includes average, sum, median, and maximum latency together with average, median, and maximum speedup for every set. A progress bar is written to standard error. Text, JSON, CSV and XLSX output include the input shapes, parameter values and speedups. Input shapes preserve their input names and dataset grouping. The comparison uses temporary worker caches and never modifies the machine tuning cache. The selected set is reported but not used by later kernels. See Tune kernel thresholds for criteria, timeout handling, result locations, cache removal, and the Python analysis API.