onnx_light.onnx_core.quantization#
See the quantization guide for the profile catalogue, parameter reference and numerical contracts, and the runnable Python tutorial for examples covering every profile.
The usual NumPy path is numpy_helper.from_array ->
quantize_tensor_proto -> EncodedValueProto ->
dequantize_tensor_proto -> numpy_helper.to_array.
Profiles define portable storage defaults, not calibration algorithms or
vendor-compatible binary formats.
The explicit ORT_MATMULNBITS_INT2/INT4/INT8 profiles are the exception:
make_matmul_nbits_plan and export_matmul_nbits_inputs produce the
standard ONNX Runtime operator inputs, not internal kernel prepacking.
For model-level shared parameters,
add_quantization_parameters declares a fixed set and
quantize_tensor_shared returns a resource-retaining SharedQuantizedValue.
materialize_quantized_value exports a self-contained message when the
declaring model will not accompany the encoded value.
Converts tensors to portable, self-describing quantized values.
Portable profiles describe onnx-light representations, not vendor binary layouts. ORT_MATMULNBITS_INT2/INT4/INT8 produce ONNX Runtime operator inputs, not internal kernel prepacking. No profile performs calibration. See Quantizes tensors into encoded values.
- class onnx_light.onnx_core.quantization.QuantizationFormat(*values)#
- class onnx_light.onnx_core.quantization.QuantizationMethod(*values)#
- class onnx_light.onnx_core.quantization.Shape(*args, **kwargs)#
A concrete runtime shape with at most 16 integer dimensions. An empty shape represents a scalar.
- onnx_light.onnx_core.quantization.add_quantization_parameters(model, name, storage_type, logical_type, *, scales, zero_points=None, offsets=None, codebooks=None, permutation=None, forward=None, inverse=None, outliers=None)#
Adds a fixed numerical parameter set backed by root graph initializers.
Numerical arguments accept TensorProto or NumPy arrays. Descriptors and tensors are copied; the model is changed only after successful validation. :returns: The compact StructTypeProto used by encoded values referring to this set.