Unary Elementwise Performance Roadmap#
- Date:
2026-08
- Updated:
2026-09-02
discussion
Objective#
The objective is to provide one prepared CPU engine for every ONNX unary
elementwise operator, with portable scalar semantics and tuned SIMD kernels for
x86 and ARM. The priority corpus must reach at least 1.0x ONNX Runtime
median performance, with no priority case below 0.9x. Correctness includes
all supported data types, attributes, special values, empty tensors, tails,
aliasing rules, and opset-specific behavior.
onnx-light-cpu already registers optimized kernels for Abs, Exp,
Log, and Not. These operators exist today; the roadmap extends their
shared architecture and adds the remaining operators rather than replacing
their working entry points. The common unary plan should retain their kernels
while moving selection, traversal, type conversion, accuracy policy, and
scheduling policy into shared preparation.
#556 added the direct ONNX Runtime differential matrix for the existing optimized unary kernels, including boundary semantics, vector tails, and exact execution-path checks. It strengthens the acceptance baseline but does not implement the pending operator families below, so this roadmap remains in discussion.
Scope#
This roadmap covers operators whose output elements are independently derived from one input element. Attributes and optional scalar parameters may affect the calculation, but no output element depends on another input position.
Family |
Operators |
Current status |
|---|---|---|
Basic arithmetic |
|
|
Exponential and error |
|
|
Trigonometric |
|
Pending |
Hyperbolic |
|
Pending |
Activations |
|
Pending |
Predicates and logical |
|
|
Bit and type transforms |
|
Pending |
String predicates |
|
Pending specialized path |
Parameterized unary |
|
Pending |
Operators with one required input but cross-element behavior are not unary
elementwise kernels and do not use this engine. This explicitly excludes
reductions and index selection (ArgMax, ArgMin, NonZero,
Unique), normalization and axis reductions (Softmax, LogSoftmax,
Hardmax, LpNormalization), pooling, layout transforms, random
generation, sequence/control-flow operators, non-elementwise string
operators, and shape operators. They require separate roadmaps rather than
hidden special cases here.
Unary execution plan#
Static nodes should construct an immutable UnaryElementwisePlan from the
operator, opset, input/output types, attributes, tensor size, CPU features,
thread limit, and accuracy policy. Dynamic shapes may cache plans by those
properties and element count.
The plan records:
the typed scalar fallback and selected ISA function;
input and output element sizes and whether in-place execution is legal;
normalized activation attributes or
Clipbounds;exact, correctly-rounded, or bounded-approximation semantics;
vector width, unroll factor, tail strategy, task size, and useful threads;
conversion/compute type for FP16, BF16, Float8, and integer inputs;
guards that reject incompatible shapes, types, attributes, or opsets.
Traversal and calculation remain separate. A hot loop calls a typed function selected once by the plan; it must not branch on the operator or perform type-erased calls for every element.
Kernel families#
Native instruction kernels#
Abs, Neg, Relu, Sign, rounding, predicates, logical/bitwise
operations, and many casts map directly to SIMD arithmetic, masks, or
conversion instructions. These kernels should provide:
scalar, SSE2/AVX2, AVX-512, NEON, and SVE/SVE2 implementations;
vector tails through masks where profitable and one shared scalar fallback;
byte-valued ONNX
BOOLoutput without bit packing;defined integer behavior without relying on C++ signed-overflow assumptions;
alias-safe in-place execution only where input and output representation permit it.
Transcendental kernels#
Trigonometric, hyperbolic, exponential, logarithmic, error, sigmoid, and GELU families require range reduction and polynomial or rational approximations. Share primitives instead of independently approximating each operator:
Expfeeds sigmoid, softplus, swish, mish, and parts of hyperbolic functions;Log/Log1pfeed softplus and inverse hyperbolic functions;shared sine/cosine range reduction computes
SinandCostogether;Erffeeds the exact GELU formulation;reciprocal and reciprocal-square-root helpers retain explicit refinement and error contracts.
Every approximation documents maximum ULP or relative error over normal, subnormal, overflow, and underflow ranges. NaN payload behavior, infinities, domain errors, and signed zero must match the ONNX/host reference contract. Fast-math behavior is opt-in and never silently replaces the default kernel.
Composite activations#
Activation kernels should compose shared vector primitives but remain fused in
one traversal. For example, Mish must not allocate outputs for softplus and
tanh, and HardSwish should keep its clamp and multiply in registers.
Attribute-bearing activations normalize constants in the plan so the hot loop
loads broadcast vectors rather than parsing attributes.
Types#
Type family |
Preferred implementation |
Required behavior |
|---|---|---|
FP32/FP64 |
Native SIMD or documented vector approximation |
Preserve special values, signed zero, domain, and accuracy contracts. |
FP16/BF16 |
Native arithmetic or vector conversion to FP32 |
Convert vectors, compute safely, and narrow once; avoid scalar decode/encode loops. |
Float8 |
Explicit vector decode, FP16/FP32 compute, explicit encode |
Keep each ONNX Float8 encoding and saturation rule distinct. |
Integers |
Native vector arithmetic/masks where defined |
Match operator type constraints and exact overflow semantics. |
|
Byte-vector masks |
Preserve the runtime byte representation. |
Cast output |
Typed conversion kernels |
Match truncation, saturation, string, and unsupported-conversion rules for the selected opset; non-numeric conversions may use a specialized fallback outside the SIMD numeric loop. |
String tensors |
Exact specialized loops |
|
Parallel scheduling#
Unary arithmetic is usually memory-bandwidth-bound, while transcendental
operators are compute-bound. Registered kernels already execute through the
onnx-light session executor, with no private onnx-light-cpu scheduler.
Prepared kernels therefore select measured limits for the session-owned
executor by operator family, type, processor profile, and ISA:
cheap kernels stay single-threaded until enough bytes amortize dispatch;
expensive functions may parallelize at much smaller element counts;
blocks begin on cache-line and SIMD boundaries;
thread count is capped when memory bandwidth or available blocks saturate;
caller-owned pools and the internal pool must not oversubscribe each other.
Current tuning contract#
Abs registers tuning ABI 4, while Exp/Log and Not register
tuning ABI 2 for every supported exact operator/input-type key with
implementation="simd_dispatch". The shared parameters are:
parallel.threshold_bytes: minimum input bytes before executor dispatch; zero disables dispatch;parallel.target_block_bytes: target input bytes per task and always positive;parallel.max_participants: executor participant ceiling; zero uses every participant exposed by the session;parallel.cost_model: one delegates grain and participant selection to the session executor, while zero retains the fixed calibrated thresholds.
For FP32 Abs, Exp, and Log, the executor cost follows the SIMD
implementation selected at runtime. The existing AVX2-calibrated costs remain
unchanged; implementations scale them by 8 / SIMD lanes. As the
memory-bound exception, Abs is not reduced below 1.0 for AVX-512 because
doing so causes the shared model to select ranges that are too small. INT32
Abs uses the same read/write cost and a compute cost of 1.0. Processor
profiles can continue to refine thresholds, block sizes, and participant
limits without changing this portable fallback.
Abs and Not additionally accept parallel.preferred_participants.
Zero leaves the participant count automatic. A positive value requests that
exact count once parallel execution is worthwhile, but the executor still
clamps it to parallel.max_participants and the session limit.
Abs also accepts memory.streaming_store_threshold_bytes; zero disables
non-temporal stores, while a positive value enables them for FP32 tensors at
or above that total input size. The portable default is zero: streaming stores
remain opt-in because their benefit depends on worker range size and memory
topology, and every worker must publish them with its own fence.
The portable defaults retain SIMD execution inline through the measured small/medium-tensor region:
Operator and input types |
Threshold |
Target block |
Maximum participants |
|---|---|---|---|
|
2 MiB |
256 KiB |
32 |
|
512 KiB |
256 KiB |
32 |
|
2 MiB |
512 KiB |
executor maximum |
|
4 MiB |
128 KiB |
executor maximum |
|
8 MiB |
64 KiB |
executor maximum |
|
2 MiB |
256 KiB |
32 |
|
1 MiB |
128 KiB |
32 |
|
64 KiB |
1,600 bytes |
executor maximum |
|
512 KiB |
64 KiB |
32 |
|
2 MiB |
256 KiB |
32 |
Processor profiles may override these values without changing kernel code or introducing process-global mutable scheduling state.
Fusion#
Unary kernels expose typed functions and plan metadata to the shared
ElementwisePlan described by the binary roadmap. Fusion may combine unary
and binary nodes only when evaluation order, broadcasting, types, aliasing,
and graph lifetimes permit it. Standalone unary parity does not depend on
fusion, but fusion is required to remove intermediate tensor traffic in common
activation and normalization expressions.
Benchmark corpus#
The corpus compares isolated kernels and representative fused expressions against ONNX Runtime. It covers:
sizes from zero and scalar tensors through bandwidth-saturating tensors;
every SIMD tail and deliberately misaligned input/output addresses;
every supported type, attribute boundary, and opset behavior;
NaN, infinities, signed zero, subnormals, domain boundaries, and saturation;
single-thread, physical-core, and logical-core configurations;
latency, throughput, effective bandwidth, selected ISA, and thread count;
accuracy histograms and worst-case error for approximate functions.
Shared CI enforces correctness. Tight performance and numerical-search gates run on pinned machines and preserve raw samples and environment metadata.
Completed foundations#
The dedicated Exp and Log ONNX Runtime Parity Roadmap is complete through
onnx-light-cpu #315. Its numerical gates,
AVX2+FMA and AVX-512 kernels, benchmark corpus, operator-specific scheduling,
and preserved evidence are inputs to this roadmap. Unary PR03 reuses those
Exp/Log implementations and primitives; it does not reimplement their
parity work.
The Runtime Execution Controls Roadmap is also complete through onnx-light-cpu #271 and #314. Registered kernels use the session-owned executor, and the existing onnx-light processor-aware tuning registry supplies the profile-resolution foundation. The remaining unary PRs build on these completed foundations; they do not add a private scheduler.
Remaining pull-request sequence#
The following table is the single source of truth for the unary roadmap.
PR |
Scope |
Merge criterion |
Depends on |
Status |
|---|---|---|---|---|
Unary PR01 |
Corpus, plan, registry, and scalar semantics. |
|
Completed runtime foundation |
Pending |
Unary PR02 |
Native arithmetic, predicates, bits, and numeric casts. |
Basic arithmetic, rounding, sign, Relu-family native operations,
|
PR01 |
Pending |
Unary PR03 |
Reuse Exp/Log primitives for reciprocal, sqrt, and composite activations. |
The merged |
PR01, PR02 |
Pending |
Unary PR04 |
Trigonometric, hyperbolic, inverse, and error functions. |
Shared range reduction and approximation primitives implement sin/cos/tan, inverse trigonometric, hyperbolic, inverse hyperbolic, and erf on x86 and ARM within the documented numerical limits. |
PR03 |
Pending |
Unary PR05 |
Low-precision and remaining conversion families. |
FP16/BF16/Float8 vector conversion or native kernels cover all
applicable operators; non-numeric casts and |
PR02 through PR04 |
Pending |
Unary PR06 |
Session-executor tuning and fusion integration. |
Processor-aware limits submitted to the session executor scale
compute-bound kernels and cap bandwidth-bound kernels without
small-tensor regressions or a private scheduler. Unary functions
integrate with |
PR02 through PR05; Binary PR07 |
Pending |
Unary PR07 |
Final correctness and parity gate. |
Every in-scope operator/type passes differential tests; median priority performance is at least 1.0x ONNX Runtime with no priority case below 0.9x. This PR remains open while any target fails. |
PR01 through PR06 |
Pending |
Unary PR07 is the final unary roadmap PR.