Note
Go to the end to download the full example code.
Benchmark Abs: onnxruntime vs onnx-light + onnx-light-cpu#
This example compares three ways of computing the elementwise absolute value
of a float32 array across a range of input sizes:
onnxruntime - running a single-node
AbsONNX model.onnx-light + onnx-light-cpu - the SIMD-accelerated
Abskernel thatonnx-lightdispatches to. The same ONNX model used by onnxruntime is evaluated by anonnx-lightReferenceEvaluatoron which theonnx-light-cpuAbskernel has been registered (onnx_light_cpu.register_kernels()); the kernel provides runtime AVX-512/AVX2/AVX/SSE2 dispatch.numpy -
numpy.abs(), used as a reference baseline.
The three back-ends compute the same result; the goal here is to see how their timings evolve as the array grows from a few hundred to a hundred million elements.
Setup#
Import the onnx-light-cpu registration helper and report which SIMD level
the current CPU provides. The mapping is 0=None, 1=SSE2, 2=AVX,
3=AVX2 and 4=AVX512. onnx-light is not published on PyPI, so when
it is not importable the example falls back to calling the compiled kernel
directly; this keeps the documentation build working everywhere.
import importlib.util
import time
import numpy as np
import onnx
import onnxruntime
from onnx import TensorProto, helper
from onnx_light_cpu import register_kernels
from onnx_light_cpu.onnx_py._cpukernels import (
abs as cpu_abs,
detect_simd_level,
has_cpu_kernels,
)
_HAS_ONNX_LIGHT = importlib.util.find_spec("onnx_light") is not None
_SIMD_NAMES = {0: "scalar", 1: "SSE2", 2: "AVX", 3: "AVX2", 4: "AVX-512"}
assert has_cpu_kernels()
level = detect_simd_level()
simd_name = _SIMD_NAMES.get(level, level)
print(f"CPU kernels available, SIMD level: {level} ({simd_name})")
CPU kernels available, SIMD level: 4 (AVX-512)
Build the onnx-light evaluator#
onnx-light evaluates the same model with its C++ runtime. Registering the
onnx-light-cpu kernels overrides the built-in Abs so every Abs node in
the model dispatches to the SIMD-accelerated kernel. When onnx-light is
not installed the same kernel is invoked directly through the compiled
extension so the benchmark still runs.
if _HAS_ONNX_LIGHT:
from onnx_light.onnx.reference import ReferenceEvaluator
light_session = ReferenceEvaluator(model)
register_kernels(light_session)
def run_light(inp):
return light_session.run(None, {"X": inp})[0]
light_label = "onnx-light + onnx-light-cpu"
else:
def run_light(inp):
return cpu_abs(inp)
light_label = "onnx-light-cpu"
Timing helper#
Each candidate is called repeat times and the best (minimum) wall-clock
time is kept to reduce the impact of scheduling noise. The number of repeats
shrinks as the arrays grow so the whole benchmark stays fast.
def measure(func, repeat):
best = float("inf")
for _ in range(repeat):
start = time.perf_counter()
func()
best = min(best, time.perf_counter() - start)
return best
Run the benchmark#
For every size the same input is fed to the three back-ends. The results are
checked against numpy.abs() to make sure every implementation agrees.
sizes = [10**k for k in range(2, 9)]
rng = np.random.default_rng(0)
rows = []
for size in sizes:
inp = rng.uniform(-100.0, 100.0, size=size).astype(np.float32)
expected = np.abs(inp)
repeat = max(3, min(200, 2_000_000 // size))
numpy_time = measure(lambda inp=inp: np.abs(inp), repeat)
cpu_time = measure(lambda inp=inp: run_light(inp), repeat)
assert np.array_equal(run_light(inp), expected), size
ort_time = measure(lambda inp=inp: session.run(None, {"X": inp}), repeat)
assert np.array_equal(session.run(None, {"X": inp})[0], expected), size
rows.append((size, numpy_time, cpu_time, ort_time))
print(
f"size={size:>9} | numpy={numpy_time * 1e6:10.2f} us | "
f"onnx-light-cpu={cpu_time * 1e6:10.2f} us | "
f"onnxruntime={ort_time * 1e6:10.2f} us"
)
sizes = np.array([r[0] for r in rows])
numpy_times = np.array([r[1] for r in rows])
cpu_times = np.array([r[2] for r in rows])
ort_times = np.array([r[3] for r in rows])
size= 100 | numpy= 0.80 us | onnx-light-cpu= 1.23 us | onnxruntime= 5.57 us
size= 1000 | numpy= 1.11 us | onnx-light-cpu= 1.37 us | onnxruntime= 5.68 us
size= 10000 | numpy= 1.96 us | onnx-light-cpu= 2.03 us | onnxruntime= 6.74 us
size= 100000 | numpy= 10.92 us | onnx-light-cpu= 8.99 us | onnxruntime= 10.54 us
size= 1000000 | numpy= 104.09 us | onnx-light-cpu= 104.28 us | onnxruntime= 58.51 us
size= 10000000 | numpy= 3543.24 us | onnx-light-cpu= 3420.85 us | onnxruntime= 1391.11 us
size=100000000 | numpy= 32434.21 us | onnx-light-cpu= 30929.90 us | onnxruntime= 19875.90 us
Plot the timings#
The left panel shows the raw execution time versus the array size on a log-log scale. The right panel shows the throughput (elements processed per second), which highlights the fixed per-call overhead at small sizes.
import matplotlib.pyplot as plt
fig, (ax_time, ax_tput) = plt.subplots(1, 2, figsize=(11, 4.5))
ax_time.plot(sizes, numpy_times * 1e6, "o--", label="numpy", color="#9b7ec8")
ax_time.plot(
sizes,
cpu_times * 1e6,
"o-",
label=light_label,
color="#4a9eff",
)
ax_time.plot(sizes, ort_times * 1e6, "o-", label="onnxruntime", color="#f4a259")
ax_time.set_xscale("log")
ax_time.set_yscale("log")
ax_time.set_xlabel("array size (elements)")
ax_time.set_ylabel("time (microseconds)")
ax_time.set_title(f"Abs execution time (SIMD: {simd_name})")
ax_time.legend()
ax_tput.plot(sizes, sizes / numpy_times, "o--", label="numpy", color="#9b7ec8")
ax_tput.plot(
sizes,
sizes / cpu_times,
"o-",
label=light_label,
color="#4a9eff",
)
ax_tput.plot(sizes, sizes / ort_times, "o-", label="onnxruntime", color="#f4a259")
ax_tput.set_xscale("log")
ax_tput.set_yscale("log")
ax_tput.set_xlabel("array size (elements)")
ax_tput.set_ylabel("throughput (elements / second)")
ax_tput.set_title("Abs throughput")
ax_tput.legend()
fig.tight_layout()
plt.show()

Total running time of the script: (0 minutes 1.482 seconds)