Thread-pool dispatch in onnx-light, OpenMP, and ONNX Runtime#
This page compares how three CPU execution systems run a dependent sequence of ONNX operators:
The comparison assumes that graph optimization has not fused the operators and
that every kernel uses the execution system named in its column. The second
operator cannot start before the first has produced Y1, and the final
Gemm cannot start before Softmax has produced P. Thread-pool design
therefore changes dispatch and synchronization cost, but does not remove these
data dependencies.
Scope and terminology#
A participant is either the thread that called the parallel region or one worker. A dispatch publishes work to workers. A join below means waiting for the region, not destroying and joining the underlying operating-system threads. Persistent workers normally survive all three operators.
OpenMP is a specification rather than one runtime implementation. The OpenMP
column describes the common fork-join design used by LLVM libomp and GCC
libgomp; exact queues, barriers, spin budgets, and operating-system wait
primitives differ between those runtimes and platforms.
Overview#
Aspect |
|
OpenMP |
ONNX Runtime |
|---|---|---|---|
Pool ownership |
A session lazily leases a |
The OpenMP runtime owns teams of workers. Implementations normally retain a hot team or reusable workers after the first parallel region. |
An inference session owns an intra-operator pool by default. Applications can instead configure environment-level global pools shared by sessions. |
Worker creation |
The first lease constructs |
The first relevant parallel region initializes or expands a team, commonly using POSIX or native platform threads. Later regions reuse it. |
Pool creation creates the extra intra-operator workers. The default participant count targets physical cores and includes the calling thread. |
Work publication |
A short mutex-protected update marks the generation odd, installs the callable and block count, notifies each admitted worker, and publishes an even generation. Spinning workers acquire work without the mutex. |
The compiler outlines the region. |
ORT submits ranges or tasks to its modified Eigen non-blocking pool. Workers can consume assigned work and attempt to steal available work. |
Work assignment |
Static: block 0 runs on the caller and block |
|
Operator helpers choose a task count from work cost. Scheduling is more
dynamic than |
Idle waiting |
Workers spin for the resolved iteration or duration budget and then
park on a |
Workers normally spin or yield for a runtime-defined interval and then use an implementation- and platform-specific blocking wait. |
Spinning is enabled by default. ORT supports disabling it or selecting a calibrated duration and exponential pause backoff before workers sleep. |
Region completion |
Workers decrement an atomic remaining count. The caller spins, then waits on a second condition variable if work is still outstanding. |
A fork-join barrier, commonly generation- or sense-based and optimized with atomics or a tree, releases the primary thread after the team arrives. |
The caller waits for the submitted intra-operator work. Pool barriers and task counters complete the operator before graph execution advances. |
Nested work |
Nested regions run inline by default. |
Controlled by OpenMP nesting and active-level settings; a nested region may serialize or form another team. |
The thread pool detects parallel sections and limits nested parallelism; operator implementations also use cost thresholds. |
Concurrent callers |
Dispatch metadata is serialized per shared executor. Compatible sessions may share workers, but their parallel regions do not execute concurrently on that executor. |
Behavior depends on the runtime and team configuration. Independent host threads can request teams and may contend or oversubscribe. |
Per-session pools can execute independently and may oversubscribe the machine. Global pools trade isolation for shared capacity. |
Shutdown |
Releasing the final compatible lease sets an atomic stop flag, changes
the generation, notifies all workers, and joins every |
Workers usually live until runtime or process teardown; implementation shutdown joins or releases the native threads. |
A per-session pool ends with its session; a global pool ends with its environment. Custom thread callbacks expose creation and joining. |
The concrete three-operator execution#
The following timeline shows the synchronization visible to the graph
executor. D is dispatch, B is the region-completion barrier, and
idle means either spinning or parked:
graph caller: D Gemm 1 B | D Softmax B | D Gemm 2 B
workers: wake/work/idle | wake/work/idle | wake/work/idle
dependency: Y1 | P | Y2
onnx-light#
RuntimeSession::Run acquires the executor on first use and installs it on
the calling thread and RuntimeContext. Each parallel kernel calls
CpuExecutor::ParallelFor:
The kernel may stay inline when its work is below
grain. Otherwise,ParallelFordivides the range into at most one block per admitted participant.ThreadPool::Runlocks the executor’s region and state mutexes, marksgeneration_odd, installs one callable, and notifies admitted workers. It publishes an even generation immediately before unlocking.The caller computes block 0 while worker
j - 1computes blockj.Workers decrement
remaining_. The caller first spins and then waits oncv_done_if needed.The workers spin for the next generation and eventually park on their individual condition variables. The nearby
Softmaxor secondGemmcan reach them while they are still spinning and avoid an operating-system wake-up.
The mutexes do not cover matrix multiplication or softmax arithmetic. They
only serialize publication and concurrent regions. Workers without an assigned
block are not notified. Spinning workers accept an even generation only when
it remains unchanged across an acquire load of the selected-worker count.
The count’s release store follows the odd-generation marker, so a worker
cannot combine an old generation with a newer count. Once selected, its
outstanding contribution to remaining_ prevents replacement of the
callable until it finishes. This avoids serializing warm workers on the state
mutex without allowing concurrent reads of a changing payload.
Parked workers still check the generation predicate under the state mutex, so notifying before publication does not lose wakeups.
OpenMP#
For three kernels implemented as three OpenMP parallel loops, the compiler outlines each loop body and emits three runtime fork-join calls:
The first call obtains a team and may pay lazy worker creation.
A static
Gemmloop gives each team member a deterministic tile range.The join barrier makes
Y1visible before the primary thread enters theSoftmaxregion.The same fork-join sequence occurs for
Softmaxand the secondGemm. Workers are normally reused rather than recreated.
Static scheduling avoids a lock per tile. Dynamic scheduling can balance irregular tiles but adds atomic scheduling traffic. Keeping one outer OpenMP region around all three operators could remove two team releases, but it would require team-aware kernels and explicit barriers between dependent operators. Calling independently parallel kernels from that outer region can instead trigger nested parallelism or serialization, so inference runtimes generally dispatch each operator separately.
ONNX Runtime#
With the default sequential graph execution mode, the three nodes run in graph order and use the session’s intra-operator pool:
The first
Gemmasks MLAS and ORT’s thread-pool helpers to parallelize suitable matrix tiles.The operator waits for its intra-operator tasks before the graph executor dispatches
Softmax.Softmaxparallelizes only when its row count and estimated work justify pool overhead; otherwise it can run on the caller while the extra workers remain idle.The second
Gemmreuses the same persistent intra-operator workers.
ORT_PARALLEL adds a distinct inter-operator pool for independent graph
branches. It does not overlap this linear sequence because each node consumes
its predecessor’s output. ORT’s default worker spinning favors the short gaps
between these operators; disabling spinning saves CPU and power but can add a
sleep-to-running transition to the next dispatch.
Where time is spent#
For one inference after pool initialization, a useful decomposition is:
Worker creation is outside this steady-state equation. If a program constructs and destroys a session for every inference, pool creation and joining must be added and can dominate small models.
Large Gemm operators normally dwarf mutex and dispatch costs. The middle
Softmax is the important boundary case: waking a team can cost more than a
small softmax, so all three systems need a threshold that leaves insufficient
work inline. For medium work, the spin policy decides whether the next
operator observes a worker already running or pays scheduler wake-up latency.
The systems optimize different constraints:
onnx-lightfavors explicit ownership, deterministic static assignment, compatible-session sharing, and inspectable policy.OpenMP offers highly tuned fork-join barriers and several schedules, but the exact lifetime and waiting behavior belongs to the selected OpenMP runtime.
ONNX Runtime favors general operator task scheduling, work stealing, and configurable per-session or global pools.
Consequently, replacing one uncontended mutex is unlikely to change a large
Gemm–Softmax–Gemm pipeline. Measurements should separate first
run from steady state, report whether workers spin or park, and include a small
softmax case where dispatch overhead is measurable.
Implementation references#
onnx-light: CPU executor, persistent thread pool, and Session execution policies and shared CPU pools.ONNX Runtime: Thread management, thread-pool construction, and modified Eigen non-blocking pool.
OpenMP: LLVM fork-join runtime, LLVM wait and release primitives, GCC team lifecycle, and GCC barriers.