onnx_light.onnx_core.quantization#

See the quantization guide for the profile catalogue, parameter reference and numerical contracts, and the runnable Python tutorial for examples covering every profile.

The usual NumPy path is numpy_helper.from_array -> quantize_tensor_proto -> EncodedValueProto -> dequantize_tensor_proto -> numpy_helper.to_array. Profiles define portable storage defaults, not calibration algorithms or vendor-compatible binary formats. The explicit ORT_MATMULNBITS_INT2/INT4/INT8 profiles are the exception: make_matmul_nbits_plan and export_matmul_nbits_inputs produce the standard ONNX Runtime operator inputs, not internal kernel prepacking.

For model-level shared parameters, add_quantization_parameters declares a fixed set and quantize_tensor_shared returns a resource-retaining SharedQuantizedValue. materialize_quantized_value exports a self-contained message when the declaring model will not accompany the encoded value.

Converts tensors to portable, self-describing quantized values.

Portable profiles describe onnx-light representations, not vendor binary layouts. ORT_MATMULNBITS_INT2/INT4/INT8 produce ONNX Runtime operator inputs, not internal kernel prepacking. No profile performs calibration. See Quantizes tensors into encoded values.

class onnx_light.onnx_core.quantization.QuantizationFormat(*values)#
class onnx_light.onnx_core.quantization.QuantizationMethod(*values)#
class onnx_light.onnx_core.quantization.Shape(*args, **kwargs)#

A concrete runtime shape with at most 16 integer dimensions. An empty shape represents a scalar.

append(self, dim: int) → None#

Appends one integer dimension.

dims(self) → list[int]#

Returns a copy of the dimensions as a list.

empty(self) → bool#

Returns whether this is a scalar shape.

product(self) → int#

Returns the checked dimension product, or one for a scalar.

rank(self) → int#

Returns the number of dimensions.

onnx_light.onnx_core.quantization.add_quantization_parameters(model, name, storage_type, logical_type, *, scales, zero_points=None, offsets=None, codebooks=None, permutation=None, forward=None, inverse=None, outliers=None)#

Adds a fixed numerical parameter set backed by root graph initializers.

Numerical arguments accept TensorProto or NumPy arrays. Descriptors and tensors are copied; the model is changed only after successful validation. :returns: The compact StructTypeProto used by encoded values referring to this set.