Next Steps#

Date:

2026-08

Updated:

2026-10

Started#

Plan

Remaining work

Complete native kernel parallelization

Migrate every measured payload-scale kernel family to the session executor or record a benchmark-backed serial exemption; finish Gemm tuning, ARM64/x86-64 default promotion, and ORT attribution.

Using onnx-light fast loading in onnxruntime

Implement issue #4612 in an ONNX Runtime fork: retain mapped-payload owners in SessionState, use direct reads for ineligible tensors, run the four-configuration benchmark, and submit the upstream PR. All native dependencies through #4623 are complete.

Discussion#

Plan

Contribution

Proto schema inheritance

Reuses common schema fields without changing the flat wire format; independent of the prepared-value and persistent-state plan.

Model resolution before weight loading

Determines the final graph and live payloads before parallel reads or ONNX Runtime handoff.

Completed#

2025#

Plan

Contribution

Building onnx_proto, the protobuf-free ONNX schema

Supplies the protobuf-free model representation used by startup and ONNX Runtime.

2026#

Plan

Contribution

Fast-loading implementation sequence

Orders the four startup plans: bug fixes, prepared execution, native completion, then final ONNX Runtime integration.

Model-loading bug fixes

Makes parsing, external data, and initializer materialization reliable before asynchronous work starts.

Prepared and asynchronous execution

Provides the dependency graph, bounded scheduling, kernel creation, and prepacking needed by parallel startup.

Completing native fast model loading

Connects adaptive I/O, model resolution, prepared tensors, and first-token overlap before any new work in ONNX Runtime.

ParallelFor profiling and hardware counters

Measures work decomposition, utilization, and hardware counters so tuning decisions have evidence.

Compact GraphBuilder authoring and runtime walkthroughs

Supplies reproducible models and workflows used to exercise the roadmap.

Porting the ONNX C++ library on top of onnx_proto

Supplies ONNX validation and transformation without libprotobuf.

Operator kernels and the C++ backend tests

Provides the native kernels and correctness corpus required before parallel variants are accepted.

Symbolic gradients for ONNX graphs

Supports training graphs independently from the runtime roadmap.

Integrating onnx-light into onnxruntime (PR #29723)

Establishes the ONNX Runtime build-time integration extended by the optimized startup contract.

Reducing the lib_onnx_proto binary size

Keeps the library embedded by ONNX Runtime small.

Processor-aware kernel tuning

Provides schemas, defaults, calibration, user overrides, immutable snapshots, and persistent machine profiles.

Buffer-reuse arenas

Supplies ownership-aware reusable storage for startup buffers and persistent state.

Pattern-based optimization in GraphBuilder

Finalizes graph rewrites before live payload selection and ONNX Runtime handoff.

Session execution policies and shared CPU pools

Supplies the shared executor used by parallel kernels and startup tasks.

Custom, quantized, and persistent values

Supplies structured and quantized values, graph-declared zero-copy feedback, contiguous KV reuse, paged quantized caches, and end-to-end decode validation with allocation and copy measurements.

Consolidated design references#

The following proposals are retained as historical detail and format examples. Their implementation sequences are superseded by Custom, quantized, and persistent values.