PaddleOCR-VL-1.5 ONNX Optimized (Quantized)
4
6 commits
1 linked in READMEs
updated Aug 31, 2026
This repository provides the ONNX version of PaddleOCR-VL-1.5, specifically optimized for CPU inference. The models have been quantized using Intel® ONNX Neural Compressor (INC) and NNCF to significantly reduce memory usage and increase inference speed.
This project offers multiple levels of quantization so you can choose the best balance between speed and accuracy for your hardware.
The vision component handles the initial image processing and maps visual features to the language space.
vision_encoder.onnx: Full FP32 precision.vision_encoder_q8.onnx: 8-bit dynamic quantized version.vision_encoder_q4.onnx: 4-bit weight-only quantization, compressed using NNCF (Data-free).The autoregressive decoder responsible for generating the text/structured output.
decoder.onnx: Full FP32 precision language model.decoder_q8.onnx: 8-bit dynamic quantized version.decoder_q4.onnx: 4-bit quantized version using the GPTQ algorithm via the ONNX Neural Compressor.embed.onnx: The shared embedding layer required for token-to-vector conversion.Check out this colab notebook Demo , I used OpenVINO backend ,but you can used what ever backend depends on the hardware onnx providers
PaddleOCR-VL-1.5 ONNX Optimized (Quantized)
4
6 commits
1 linked in READMEs
updated Aug 31, 2026
This repository provides the ONNX version of PaddleOCR-VL-1.5, specifically optimized for CPU inference. The models have been quantized using Intel® ONNX Neural Compressor (INC) and NNCF to significantly reduce memory usage and increase inference speed.
This project offers multiple levels of quantization so you can choose the best balance between speed and accuracy for your hardware.
The vision component handles the initial image processing and maps visual features to the language space.
vision_encoder.onnx: Full FP32 precision.vision_encoder_q8.onnx: 8-bit dynamic quantized version.vision_encoder_q4.onnx: 4-bit weight-only quantization, compressed using NNCF (Data-free).The autoregressive decoder responsible for generating the text/structured output.
decoder.onnx: Full FP32 precision language model.decoder_q8.onnx: 8-bit dynamic quantized version.decoder_q4.onnx: 4-bit quantized version using the GPTQ algorithm via the ONNX Neural Compressor.embed.onnx: The shared embedding layer required for token-to-vector conversion.Check out this colab notebook Demo , I used OpenVINO backend ,but you can used what ever backend depends on the hardware onnx providers