kernel#

Lists every native kernel currently registered by the runtime:

onnx-light kernel --list

Select one or more kernels by operator name to display every registered tuning schema, including its element type, implementation, ABI, parameter names, portable defaults, and active values:

onnx-light kernel --kernel Gemm --kernel Softmax

A domain-qualified identifier such as ai.onnx:Gemm may be used when operator names overlap across domains. --library, --device, --dtype and --impl select the complete tuning-key identity. Every dimension defaults to all registered values:

onnx-light kernel --kernel Gemm \
    --library onnx_light --device CPU \
    --dtype FLOAT --impl portable

Devices accept CPU, Undefined, GPU0 through GPU8191, or their numeric value. Dtypes accept ONNX names such as FLOAT or their integer values. The output includes the library, device, dtype and implementation for every tuning schema, and the device also filters the native registrations considered by --list and --kernel. Fixed-policy kernels are reported with no tunable parameters. Add --json to either form for deterministic machine-readable output.

Add --tune to calibrate and persist every matching tuning schema for one selected native kernel. The command fails when the name or device selection resolves to zero or multiple native kernels. It prints the active parameters before and after calibration:

onnx-light kernel --kernel Gemm \
    --dtype FLOAT --impl portable \
    --tune --verbose

--verbose reports calibration progress on stderr. --cache selects an explicit cache file, while --maximum-duration-ms and --maximum-memory-mb bound each calibration; zero uses the callback default. The text output reports the selected budgets and, when parameters are persisted, the machine tuning cache path and update status. With --json, the tuning_options, before, calibrations and after sections are machine-readable.

Use --parameter NAME=default,VALUE,... to compare explicit integer values with the kernel’s calibration workload:

python -m onnx_light kernel --kernel Gemm \
    --dtype FLOAT --impl portable --tune \
    --parameter parallel.minimum_tasks=default,1,2,4

default resolves to the current active value and must come first. It is the baseline for every reported speedup. The command prints each value’s elapsed time and speedup side by side, selects the fastest value, and persists it in the same tuning cache as a regular calibration.