kernel#
Lists every native kernel currently registered by the runtime:
onnx-light kernel --list
Select one or more kernels by operator name to display every registered tuning schema, including its element type, implementation, ABI, parameter names, portable defaults, and active values:
onnx-light kernel --kernel Gemm --kernel Softmax
A domain-qualified identifier such as ai.onnx:Gemm may be used when
operator names overlap across domains. --library, --device, --dtype
and --impl select the complete tuning-key identity. Every dimension
defaults to all registered values:
onnx-light kernel --kernel Gemm \
--library onnx_light --device CPU \
--dtype FLOAT --impl portable
Devices accept CPU, Undefined, GPU0 through GPU8191, or their
numeric value. Dtypes accept ONNX names such as FLOAT or their integer
values. The output includes the library, device, dtype and implementation for
every tuning schema, and the device also filters the native registrations
considered by --list and --kernel. Fixed-policy kernels are reported
with no tunable parameters. Add --json to either form for deterministic
machine-readable output.
Add --tune to calibrate and persist every matching tuning schema for one
selected native kernel. The command fails when the name or device selection
resolves to zero or multiple native kernels. It prints the active parameters
before and after calibration:
onnx-light kernel --kernel Gemm \
--dtype FLOAT --impl portable \
--tune --verbose
--verbose reports calibration progress on stderr. --cache selects an
explicit cache file, while --maximum-duration-ms and
--maximum-memory-mb bound each calibration; zero uses the callback default.
The text output reports the selected budgets and, when parameters are
persisted, the machine tuning cache path and update status. With --json,
the tuning_options, before, calibrations and after sections are
machine-readable.
Use --parameter NAME=default,VALUE,... to compare explicit integer values
with the kernel’s calibration workload:
python -m onnx_light kernel --kernel Gemm \
--dtype FLOAT --impl portable --tune \
--parameter parallel.minimum_tasks=default,1,2,4
default resolves to the current active value and must come first. It is the
baseline for every reported speedup. The command prints each value’s elapsed
time and speedup side by side, selects the fastest value, and persists it in the
same tuning cache as a regular calibration.