Tune kernel thresholds#
onnx-light separates two kinds of tuning values:
portable defaults are compiled into the kernel library and always exist;
calibrated profiles are measured for one processor and effective thread count, then optionally persisted in a cache.
A cache profile overrides the portable defaults only when its complete tuning key, processor descriptor, and effective thread count match. It never changes the ONNX model or the numerical contract of the operator.
Inspect parameters from Python#
The tuning API is exposed by the full Python build:
from onnx_light import kernel_tuning
from onnx_light.onnx import TensorProto
report = kernel_tuning.kernel_tuning_parameters(
kernel="Gemm",
element_type=int(TensorProto.FLOAT),
)
for kernel in report["kernels"]:
print(kernel["parameter_names"])
print("portable:", kernel["defaults"])
print("local cache:", kernel["cached_values"])
print("active:", kernel["active_values"], kernel["active_source"])
Each result identifies the tuning library, kernel, implementation,
ONNX element_type, CPU device, and tuning_abi. cached_values is
None when the cache has no profile matching both the local processor and
the requested effective thread count. active_source distinguishes
portable_default, published_profile (loaded, calibrated, or explicitly
overridden), and a statically registered processor registered_profile.
Without a kernel filter, the function returns every registered exact key.
The optional library, implementation, element_type, path, and
num_threads arguments narrow the query. The report also includes the cache
path, parse status, and diagnostics.
Change local values from Python#
set_kernel_tuning_parameters accepts a partial update. Unspecified values
come from the matching cached profile when one exists, or from the portable
defaults otherwise. The complete result is validated before the cache is
modified:
update = kernel_tuning.set_kernel_tuning_parameters(
"Gemm",
int(TensorProto.FLOAT),
{
"parallel.minimum_tasks": 4,
},
)
assert update["status"] == "updated", update["diagnostics"]
By default the update is persisted atomically in the default cache and loaded
into the current process. Pass path=... for another cache,
num_threads=... for a profile scoped to that effective thread count, or
load=False to persist without activating it. Unknown names, wrong Python
types, and values rejected by the kernel schema raise an exception without
changing the file.
inspect_kernel_tuning_cache(path=None, num_threads=0) returns every
persisted profile and marks the profiles matching the local processor and
thread count with local=True. It does not change active runtime values.
default_kernel_tuning_cache_path() returns the default path.
See Inspect, change, and calibrate kernel tuning from Python for an executable gallery example that combines discovery, a temporary validated update, cache inspection, inference, and bounded calibration.
Calibrate one kernel from the command line#
Select one exact native kernel and add --tune to calibrate and persist its
parameters:
python -m onnx_light kernel \
--kernel Abs --dtype FLOAT --impl portable \
--tune
Use --cache to select another cache, --json for machine-readable output,
and --maximum-duration-ms or --maximum-memory-mb to bound calibration.
The Python functions propose_kernel_tuning_updates and
apply_kernel_tuning_updates remain available for bulk workflows.
To compare explicit values for one integer parameter, list default first
so the current active value becomes the speedup baseline:
python -m onnx_light kernel \
--kernel Abs --dtype FLOAT --impl portable --tune \
--parameter parallel.minimum_elements=default,16384,32768,65536
The existing calibration cases measure every value side by side. The fastest validated value is persisted, while text and JSON output report elapsed times, speedups, the baseline, and the selected value.
parallel.minimum_elements is the serial/parallel crossover: a loop with
exactly that many elements is eligible for parallel execution. Once the loop is
dispatched, the executor derives the block grain from the number of admitted
participants, independently from the crossover value.
Optimize over backend cases#
Use onnx-light backend when the objective is latency over a list of
backend cases rather than a kernel’s synthetic calibration workload. A regular
expression is required so the corpus is explicit:
python -m onnx_light backend \
--regex "^test_cc_not.*benchmark" --mode benchmark \
--kernel Not --dtype BOOL --impl portable \
--parameter parallel.minimum_elements=default,16384,32768,65536 \
--criterion median-speedup --json
Repeat --parameter when a schema exposes interacting parameters. The command
evaluates their Cartesian product, with a limit of 256 sets. Each specification
starts with default; the resulting all-default set is the baseline.
Parameter names must be unique, integer values must be positive, and
--kernel, --dtype, and --impl must identify exactly one schema.
--criterion is required and accepts:
average,sum,median, ormax-latency, which minimize the corresponding latency across selected cases;average-speedup,median-speedup, ormax-speedup, which maximize per-case speedup relative to the all-default baseline.
Every parameter set reports all seven metrics, its timeout count, and whether
it was selected. A timed-out set has unavailable metrics. When the baseline
times out, latency criteria can still select a complete candidate, but speedup
metrics and speedup-based selection are unavailable. Progress is written to
standard error, so --json on standard output remains machine-readable:
[backend tune] [##########----------] 2/4
[backend tune] [####################] 4/4
The comparison uses temporary cache files and does not modify the machine
tuning cache. Its result is printed to standard output, returned as JSON with
--json, or written as CSV/XLSX with --output. The selected set is
advisory: kernels do not use it after the command ends. Use
set_kernel_tuning_parameters separately to persist and publish a selected
set.
Analyze measurements from Python#
The same native C++ metric analyzer is exposed in Python for measurements collected by another harness. Rows represent parameter sets, columns represent the same ordered cases, and the first row is the speedup baseline:
report = kernel_tuning.analyze_kernel_tuning_latencies(
[
[0.002, 0.008, 0.010],
[0.001, 0.004, 0.020],
[0.004, 0.004, 0.005],
],
"average-speedup",
)
print(report["selected_index"])
for metrics in report["values"]:
print(metrics)
Use None for a missing case measurement. The corresponding row is
incomplete. The return value contains criterion, selected_index, and
values. Every complete value contains average, sum, median,
average_speedup, median_speedup, max_speedup, and max_latency.
Calibrate one kernel from Python#
The built-in calibration callbacks cover Abs, Add, Gemm, Log,
Not, Sigmoid, and Tanh. The Python extension registers them when
imported. Select a kernel and optionally one or more ONNX element types:
calibration = kernel_tuning.calibrate_kernel_tuning(
"Abs",
element_types=[int(TensorProto.FLOAT)],
maximum_duration_ms=1000,
maximum_memory_bytes=128 << 20,
)
print(calibration["calibrated"])
print(calibration["diagnostics"])
print(calibration["cache_update"])
Calibration generates deterministic inputs, checks every candidate output
against the forced serial implementation, warms the implementations, and uses
median timings. The shared unary/binary crossover search requires the
configured speedup for consecutive problem sizes. Resource limits bound the
search. Inspect diagnostics to see the selected value. Schema-only keys
without a callback appear in unsupported.
Groups whose reference and candidate use the same execution path are skipped. The selected threshold is the first size in a stable winning sequence, not an extrapolation below that size. Diagnostics distinguish a crossover bracketed by measured losing and winning sizes from one at or below the smallest measured size. Gemm tries geometrically growing task counts up to 4096, stopping after stable wins or at the resource limits; its single-task inline case is excluded. When no stable win is found within those limits, the portable threshold is kept.
The selected profile is published in the current process immediately.
With the default save=True, it is also validated, locked, merged, and
atomically persisted for later processes. Use save=False for an in-memory
calibration or only_missing=True to skip an already active local profile.
Add calibration to another kernel#
A registered tuning schema does not imply that a calibration callback exists.
CalibrateRegisteredKernels reports schema-only keys in
CalibrationBatchReport::unsupported. To make another kernel calibratable:
Define a
KernelCalibrationFunctionnear the kernel implementation.Construct a
KernelCalibrationBenchmarkwith its portable parameters, deterministic cases, reference runner, candidate runner, and output validation. Setsame_execution_pathwhen different parameter values can resolve to the same path; any such case excludes its entire crossover group.Call
CalibrateKernelBenchmarkfrom that function.Register it for every supported exact key with
RegisterKernelCalibrationFunctionin the kernel’sRegisterTuningSchemasfunction.
onnx_light/onnx_extensions/kernels/kernels/math/kernel_abs.cc is the
unary example. kernel_add.cc demonstrates equal-shape and broadcasting
binary cases. A kernel with several interacting parameters, such as Gemm,
needs a kernel-specific search rather than treating every value as an
independent scalar crossover.
Promote a threshold to a compiled default#
A cache result is processor-specific. Measure several representative machines and thread counts before making it the portable value used by every machine. Keep the conservative value when crossover measurements overlap.
For kernels using ParallelTuning:
change the
portable_minimum_elementspassed toRegisterParallelTuningSchemas;change the kernel object’s initial fallback to the same value;
change
benchmark.portable_parametersin its calibration callback;add or update tests that exercise serial and parallel boundaries.
These values are in the corresponding implementation under
onnx_light/onnx_extensions/kernels/kernels/. For example, all three Abs
fallback occurrences are in kernels/math/kernel_abs.cc.
For Gemm, the compiled values are the fields of GemmTuning in
onnx_light/onnx_extensions/kernels/tuning/portable_gemm_tuning.h.
MakeGemmDefaults registers those fields as the schema defaults.
Increment tuning_abi when persisted profiles become structurally
incompatible, such as after renaming a parameter or changing its meaning or
type. A value-only default adjustment does not require an ABI change.
Locate and inspect the cache#
default_kernel_tuning_cache_path() in Python and
DefaultKernelTuningCachePath() in C++ return the exact default path:
Windows:
%LOCALAPPDATA%\onnx-light\kernel_tuning.cache;other platforms with
XDG_CACHE_HOME:$XDG_CACHE_HOME/onnx-light/kernel_tuning.cache;otherwise with
HOME:$HOME/.cache/onnx-light/kernel_tuning.cache;without any supported cache-directory environment variable:
onnx-light-kernel-tuning.cachein the current directory.
The cache is a versioned text file beginning with
onnx_light_kernel_tuning_cache 1. It can be inspected as text, but should
be modified through UpdateKernelTuningCache so validation, locking, merging,
and atomic replacement remain effective. Set KernelTuningCacheOptions::path
to use an explicit location.
Remove cached results#
Remove the default cache, or the explicit cache passed to tuning operations, through the Python API:
removal = kernel_tuning.remove_kernel_tuning_cache()
print(removal["path"], removal["removed"], removal["diagnostics"])
# Removes an explicitly selected cache instead.
removal = kernel_tuning.remove_kernel_tuning_cache("/path/to/kernel_tuning.cache")
The function takes the cache’s inter-process lock before deleting the file.
removed is False without diagnostics when the file was already absent.
After removal, a new process falls back to registered processor profiles or
portable defaults.
Removing a cache does not reconfigure kernels already initialized in the
current process. It also does not retract profiles already published into that
process’s immutable tuning-registry generations. Restart the process after
removal when subsequent sessions must stop using a profile that was previously
loaded. The equivalent native operation is
RemoveKernelTuningCache().
Load and use cached values#
Importing onnx_light.onnx_py._onnxpykernels registers the built-in tuning
schemas and automatically loads compatible profiles from the default cache.
Therefore Python RuntimeSession and ReferenceEvaluator instances use
the local default cache without another call.
An explicit path is not loaded automatically. Load it before the first session initializes its kernels:
load = kernel_tuning.load_kernel_tuning_cache(
path="/path/to/kernel_tuning.cache",
)
assert load["status"] == "loaded", load["diagnostics"]
The C++ API remains explicit for every path:
onnx_light::onnx_kernels::RegisterKernelFunctions();
rt::KernelCalibrationSelection selection;
selection.library = "onnx_light";
const rt::KernelTuningCacheLoadReport load =
rt::LoadKernelTuningCache(selection);
// Construct RuntimeSession only after loading the cache.
rt::RuntimeSession session(plan);
Check load.status, loaded, incompatible, stale, invalid,
missing, and diagnostics rather than assuming that a present file
matched. A missing, unreadable, malformed, stale, or processor-incompatible
entry leaves the compiled portable default active.
At initialization, RuntimeSession captures one immutable registry
generation, resolves a profile using its effective thread count, and copies the
typed values into each kernel. Later calls to Run do not access the
registry or reread the cache. Consequently, load a newer cache before creating
a new session; existing sessions deliberately retain their original values.