Ready-to-run container images for baidu/Unlimited-OCR.
The images are published to GHCR from every commit to main.
ghcr.io/say4n/unlimited-ocr-container:cpu
ghcr.io/say4n/unlimited-ocr-container:gpu
Use the GPU image when you have NVIDIA Docker support. Use the CPU image only for testing or small jobs; it is much slower.
The CPU image is published for linux/amd64 and linux/arm64, so it works on Intel/AMD Linux machines and Apple Silicon Docker. The GPU image is published for linux/amd64.
Apple Silicon Docker runs Linux containers, so the CPU image cannot access the macOS MPS backend or the Apple GPU. The bundled runner prefers MPS only when it is run natively on macOS with an MPS-enabled PyTorch build.
GPU:
CPU:
Both images may need Hugging Face access to baidu/Unlimited-OCR, depending on your environment.
Create folders for input files, OCR output, and model cache:
mkdir -p data outputs log models
The models mount keeps the Hugging Face cache between runs.
PDF:
docker run --rm --gpus all \
-v "$PWD/data:/data:ro" \
-v "$PWD/outputs:/workspace/outputs" \
-v "$PWD/log:/workspace/log" \
-v "$PWD/models:/models" \
ghcr.io/say4n/unlimited-ocr-container:gpu \
--pdf /data/document.pdf \
--output_dir /workspace/outputs \
--concurrency 8 \
--gpu 0 \
--image_mode gundam
Image directory:
docker run --rm --gpus all \
-v "$PWD/data:/data:ro" \
-v "$PWD/outputs:/workspace/outputs" \
-v "$PWD/log:/workspace/log" \
-v "$PWD/models:/models" \
ghcr.io/say4n/unlimited-ocr-container:gpu \
--image_dir /data/images \
--output_dir /workspace/outputs \
--concurrency 8 \
--gpu 0 \
--image_mode gundam
Markdown outputs are written to outputs/. The SGLang server log is written to log/sglang_server.log.
The CPU image supports --pdf, --image_file, and --image_dir. It does not accept GPU/SGLang options such as --gpu, --concurrency, or --server_log.
OCR text is printed to container stdout and also written as Markdown files in /workspace/outputs. The CPU image also writes visualization images with detected boxes when the upstream model returns box annotations.
PDF:
docker run --rm \
-v "$PWD/data:/data:ro" \
-v "$PWD/outputs:/workspace/outputs" \
-v "$PWD/models:/models" \
ghcr.io/say4n/unlimited-ocr-container:cpu \
--pdf /data/document.pdf \
--output_dir /workspace/outputs
Image directory:
docker run --rm \
-v "$PWD/data:/data:ro" \
-v "$PWD/outputs:/workspace/outputs" \
-v "$PWD/models:/models" \
ghcr.io/say4n/unlimited-ocr-container:cpu \
--image_dir /data/images \
--output_dir /workspace/outputs \
--image_mode gundam
Single image:
docker run --rm \
-v "$PWD/data:/data:ro" \
-v "$PWD/outputs:/workspace/outputs" \
-v "$PWD/models:/models" \
ghcr.io/say4n/unlimited-ocr-container:cpu \
--image_file /data/page.png \
--output_dir /workspace/outputs
For PDFs and single images, the stable output file is named after the input, for example outputs/document.md or outputs/page.md. Visualization files are named like outputs/page_with_boxes.jpg or outputs/document_page_0001_with_boxes.jpg. For image directories, each image gets its own Markdown and visualization files under outputs/, with nested path separators replaced by __.
If your Hugging Face setup needs a token, pass it at runtime:
docker run --rm --gpus all \
-e HF_TOKEN="$HF_TOKEN" \
-v "$PWD/data:/data:ro" \
-v "$PWD/outputs:/workspace/outputs" \
-v "$PWD/log:/workspace/log" \
-v "$PWD/models:/models" \
ghcr.io/say4n/unlimited-ocr-container:gpu \
--pdf /data/document.pdf \
--output_dir /workspace/outputs
GPU image:
--pdf PATH Convert a PDF to page images and OCR each page.
--image_dir PATH OCR every supported image under a directory.
--output_dir PATH Directory for Markdown output files.
--concurrency N Number of concurrent requests to the local SGLang server.
--gpu GPU CUDA_VISIBLE_DEVICES value inside the container.
--model_dir MODEL Hugging Face model ID or local model path.
--image_mode {gundam,base} Upstream image mode.
--server_log PATH SGLang server log path.
CPU image:
--pdf PATH Convert a PDF to page images and OCR it.
--image_file PATH OCR one image.
--image_dir PATH OCR every supported image under a directory sequentially.
--output_dir PATH Directory for Markdown output files.
--model_dir MODEL Hugging Face model ID or local model path.
--image_mode {gundam,base} Upstream single-image mode.
--max_length N Maximum generated sequence length.
--pdf_dpi N DPI used when converting PDF pages to images.
The GitHub Actions workflow at .github/workflows/publish.yml builds and pushes both images on every commit to main:
ghcr.io/say4n/unlimited-ocr-container:cpughcr.io/say4n/unlimited-ocr-container:gpughcr.io/say4n/unlimited-ocr-container:cpu-<commit-sha>ghcr.io/say4n/unlimited-ocr-container:gpu-<commit-sha>The cpu tags are multi-arch (linux/amd64, linux/arm64). The gpu tags are linux/amd64.
Local builds are only needed when changing the image definitions:
docker build --target cpu -t unlimited-ocr:cpu .
docker build --target gpu -t unlimited-ocr:gpu .
Plain docker build -t unlimited-ocr:latest . builds the CPU target.
9 commits
Python
73.1%
Dockerfile
24.9%
Shell
2.0%
Ready-to-run container images for baidu/Unlimited-OCR.
The images are published to GHCR from every commit to main.
ghcr.io/say4n/unlimited-ocr-container:cpu
ghcr.io/say4n/unlimited-ocr-container:gpu
Use the GPU image when you have NVIDIA Docker support. Use the CPU image only for testing or small jobs; it is much slower.
The CPU image is published for linux/amd64 and linux/arm64, so it works on Intel/AMD Linux machines and Apple Silicon Docker. The GPU image is published for linux/amd64.
Apple Silicon Docker runs Linux containers, so the CPU image cannot access the macOS MPS backend or the Apple GPU. The bundled runner prefers MPS only when it is run natively on macOS with an MPS-enabled PyTorch build.
GPU:
CPU:
Both images may need Hugging Face access to baidu/Unlimited-OCR, depending on your environment.
Create folders for input files, OCR output, and model cache:
mkdir -p data outputs log models
The models mount keeps the Hugging Face cache between runs.
PDF:
docker run --rm --gpus all \
-v "$PWD/data:/data:ro" \
-v "$PWD/outputs:/workspace/outputs" \
-v "$PWD/log:/workspace/log" \
-v "$PWD/models:/models" \
ghcr.io/say4n/unlimited-ocr-container:gpu \
--pdf /data/document.pdf \
--output_dir /workspace/outputs \
--concurrency 8 \
--gpu 0 \
--image_mode gundam
Image directory:
docker run --rm --gpus all \
-v "$PWD/data:/data:ro" \
-v "$PWD/outputs:/workspace/outputs" \
-v "$PWD/log:/workspace/log" \
-v "$PWD/models:/models" \
ghcr.io/say4n/unlimited-ocr-container:gpu \
--image_dir /data/images \
--output_dir /workspace/outputs \
--concurrency 8 \
--gpu 0 \
--image_mode gundam
Markdown outputs are written to outputs/. The SGLang server log is written to log/sglang_server.log.
The CPU image supports --pdf, --image_file, and --image_dir. It does not accept GPU/SGLang options such as --gpu, --concurrency, or --server_log.
OCR text is printed to container stdout and also written as Markdown files in /workspace/outputs. The CPU image also writes visualization images with detected boxes when the upstream model returns box annotations.
PDF:
docker run --rm \
-v "$PWD/data:/data:ro" \
-v "$PWD/outputs:/workspace/outputs" \
-v "$PWD/models:/models" \
ghcr.io/say4n/unlimited-ocr-container:cpu \
--pdf /data/document.pdf \
--output_dir /workspace/outputs
Image directory:
docker run --rm \
-v "$PWD/data:/data:ro" \
-v "$PWD/outputs:/workspace/outputs" \
-v "$PWD/models:/models" \
ghcr.io/say4n/unlimited-ocr-container:cpu \
--image_dir /data/images \
--output_dir /workspace/outputs \
--image_mode gundam
Single image:
docker run --rm \
-v "$PWD/data:/data:ro" \
-v "$PWD/outputs:/workspace/outputs" \
-v "$PWD/models:/models" \
ghcr.io/say4n/unlimited-ocr-container:cpu \
--image_file /data/page.png \
--output_dir /workspace/outputs
For PDFs and single images, the stable output file is named after the input, for example outputs/document.md or outputs/page.md. Visualization files are named like outputs/page_with_boxes.jpg or outputs/document_page_0001_with_boxes.jpg. For image directories, each image gets its own Markdown and visualization files under outputs/, with nested path separators replaced by __.
If your Hugging Face setup needs a token, pass it at runtime:
docker run --rm --gpus all \
-e HF_TOKEN="$HF_TOKEN" \
-v "$PWD/data:/data:ro" \
-v "$PWD/outputs:/workspace/outputs" \
-v "$PWD/log:/workspace/log" \
-v "$PWD/models:/models" \
ghcr.io/say4n/unlimited-ocr-container:gpu \
--pdf /data/document.pdf \
--output_dir /workspace/outputs
GPU image:
--pdf PATH Convert a PDF to page images and OCR each page.
--image_dir PATH OCR every supported image under a directory.
--output_dir PATH Directory for Markdown output files.
--concurrency N Number of concurrent requests to the local SGLang server.
--gpu GPU CUDA_VISIBLE_DEVICES value inside the container.
--model_dir MODEL Hugging Face model ID or local model path.
--image_mode {gundam,base} Upstream image mode.
--server_log PATH SGLang server log path.
CPU image:
--pdf PATH Convert a PDF to page images and OCR it.
--image_file PATH OCR one image.
--image_dir PATH OCR every supported image under a directory sequentially.
--output_dir PATH Directory for Markdown output files.
--model_dir MODEL Hugging Face model ID or local model path.
--image_mode {gundam,base} Upstream single-image mode.
--max_length N Maximum generated sequence length.
--pdf_dpi N DPI used when converting PDF pages to images.
The GitHub Actions workflow at .github/workflows/publish.yml builds and pushes both images on every commit to main:
ghcr.io/say4n/unlimited-ocr-container:cpughcr.io/say4n/unlimited-ocr-container:gpughcr.io/say4n/unlimited-ocr-container:cpu-<commit-sha>ghcr.io/say4n/unlimited-ocr-container:gpu-<commit-sha>The cpu tags are multi-arch (linux/amd64, linux/arm64). The gpu tags are linux/amd64.
Local builds are only needed when changing the image definitions:
docker build --target cpu -t unlimited-ocr:cpu .
docker build --target gpu -t unlimited-ocr:gpu .
Plain docker build -t unlimited-ocr:latest . builds the CPU target.
9 commits
Python
73.1%
Dockerfile
24.9%
Shell
2.0%