Complete OCR system using the DeepSeek-OCR-2 model (Released Jan 2026) with modern web interface and production-ready REST API.
โ ๏ธ IMPORTANT: This project is for DEVELOPMENT and TESTING ONLY. Not intended for production use. See LICENSE for details.
git clone https://github.com/zademy/deepSeek-ocr-docker-compose
cd deepSeek-ocr-docker-compose
cp .env.example .env
# Edit .env if needed (optional)
# Standard (no GPU reservation โ works on any Docker host)
docker compose up -d
# With NVIDIA GPU acceleration
docker compose -f docker-compose.yml -f docker-compose.gpu.yml up -d
When you first access the web interface, you'll see a "Preparar modelo" panel. Click it and wait for the download to complete (this may take several minutes depending on your internet connection).
Alternatively, use Demo Mode ("Probar sin modelo") to test the interface without downloading the model. Results in demo mode are clearly marked as simulated.
The web interface accepts JPG, PNG and WEBP up to 10 MB. PDF is not supported by the interface pipeline.
curl -X POST "http://localhost:8000/api/ocr" \
-F "file=@document.jpg" \
-F "mode=markdown"
API Documentation: http://localhost:8000/docs
Detailed documentation is available in the /docs folder:
Interactive API documentation is available at http://localhost:8000/docs when the server is running.
deepseek-ocr/
โโโ ๐ Configuration Files
โ โโโ docker-compose.yml # Docker orchestration
โ โโโ .env.example # Environment template
โ โโโ .gitignore # Git ignore rules
โ
โโโ ๐ Documentation
โ โโโ README.md # Main documentation
โ โโโ LICENSE # MIT License (Dev/Test)
โ โโโ CONTRIBUTING.md # Contribution guidelines
โ โโโ CODE_OF_CONDUCT.md # Code of conduct
โ โโโ SECURITY.md # Security policy
โ โโโ docs/ # Additional documentation
โ
โโโ ๐ Backend (FastAPI)
โ โโโ main.py # API endpoints (thin shells)
โ โโโ model_lifecycle.py # Model loading state machine + DeepSeek adapter
โ โโโ ocr_handler.py # OCR request pipeline
โ โโโ config.py # Configuration
โ โโโ Dockerfile # Container image
โ โโโ requirements.txt # Python dependencies
โ โโโ tests/ # Unit tests (pytest, no GPU needed)
โ
โโโ ๐ Frontend (HTML/JS/CSS)
โ โโโ index.html # UI structure
โ โโโ app.js # Application logic
โ โโโ styles.css # Styling
โ โโโ nginx.conf # Web server config
โ โโโ Dockerfile # Container image
โ
โโโ ๐พ Data Directories
โ โโโ uploads/ # Uploaded images
โ โโโ outputs/ # OCR results
โ
โโโ ๐งช Testing
โโโ test_api.py # API test script
Edit docker-compose.yml or create a .env file to customize:
environment:
- CUDA_VISIBLE_DEVICES=0 # GPU to use
- MODEL_NAME=deepseek-ai/DeepSeek-OCR-2
- BASE_SIZE=1024 # Base resolution
- IMAGE_SIZE=768 # Crop size (OCR-2 default)
curl -X POST "http://localhost:8000/api/ocr" \
-F "file=@image.jpg" \
-F "mode=markdown"
| Mode | Interface label (Spanish) | Recommended Use |
|---|---|---|
markdown | Documento con formato | Documents (default) |
free_ocr | Texto sin formato | General text |
grounding | Texto con posiciones | Detailed analysis |
parse_figure | Figuras y grรกficos | Charts, tables |
detailed | Descripciรณn de la imagen | Visual analysis |
The web interface shows these as "Opciones avanzadas" (collapsed by default); the API accepts the raw values.
{
"text": "# Document Title\n\nExtracted content...",
"mode": "markdown",
"processing_time": 2.5,
"image_size": [1024, 768],
"tokens": 2257
}
# Document
"<image>\n<|grounding|>Convert the document to markdown."
# General image
"<image>\n<|grounding|>OCR this image."
# No format
"<image>\nFree OCR."
# Figures
"<image>\nParse the figure."
# Detailed description
"<image>\nDescribe this image in detail."
# Start services
docker-compose up -d
# View logs
docker-compose logs -f
# Stop services
docker-compose down
# Restart
docker-compose restart
# Rebuild images
docker-compose build --no-cache
curl http://localhost:8000/health
docker-compose logs -f deepseek-ocr-api
Benchmark results with 3503ร1668 pixels image on NVIDIA A100 40GB:
| Mode | Time | Quality | Structure |
|---|---|---|---|
| Free OCR | ~24s | โญโญโญ | Basic |
| Markdown | ~39s | โญโญโญ | Complete |
| Grounding | ~58s | โญโญ | + Coords |
| Detailed | ~9s | N/A | Description |
Hardware: NVIDIA A100 40GB
# Verify NVIDIA runtime
docker run --rm --gpus all nvidia/cuda:11.8.0-base-ubuntu22.04 nvidia-smi
docker-compose logs -f deepseek-ocr-apiReduce resolution in .env:
BASE_SIZE=640
IMAGE_SIZE=512
Change ports in docker-compose.yml:
ports:
- "3001:80" # Frontend (change 3000 to 3001)
- "8001:8000" # Backend (change 8000 to 8001)
For more help, check the documentation or open an issue.
The stack runs on hosts without an NVIDIA GPU (model loads on CPU, slower):
docker compose up -d
On Apple Silicon Macs, the backend forces platform: linux/amd64 because PyTorch cu118 wheels are x86_64-only. Note that loading the 3B-parameter model on CPU needs ~8GB RAM โ increase Docker Desktop's VM memory accordingly. Use Demo Mode in the web UI to validate the interface without loading the model.
MIT License - Development and Testing Only
This software is licensed under the MIT License with specific restrictions for development and testing purposes only. It is NOT intended for production use.
โ ๏ธ Production Use Warning: If you choose to use this software in production, you do so entirely at your own risk and responsibility. The authors provide no guarantees, support, or liability for production deployments.
See the LICENSE file for full terms and conditions.
requirements.txtContributions are welcome! Please read our Contributing Guidelines before submitting PRs.
git checkout -b feature/amazing-feature)git commit -m 'feat: add amazing feature')git push origin feature/amazing-feature)Please follow our Code of Conduct in all interactions.
โ ๏ธ This project is for development and testing only.
For security concerns, please review our Security Policy.
Key Security Notes:
security labelVersion: 2.0.0
Status: Active Development
Last Updated: September 2026
Model: DeepSeek-OCR-2 (deepseek-ai)
Purpose: Development and Testing Only
If you find this project helpful, please consider:
Made with โค๏ธ for the AI community
Complete OCR system using the DeepSeek-OCR-2 model (Released Jan 2026) with modern web interface and production-ready REST API.
โ ๏ธ IMPORTANT: This project is for DEVELOPMENT and TESTING ONLY. Not intended for production use. See LICENSE for details.
git clone https://github.com/zademy/deepSeek-ocr-docker-compose
cd deepSeek-ocr-docker-compose
cp .env.example .env
# Edit .env if needed (optional)
# Standard (no GPU reservation โ works on any Docker host)
docker compose up -d
# With NVIDIA GPU acceleration
docker compose -f docker-compose.yml -f docker-compose.gpu.yml up -d
When you first access the web interface, you'll see a "Preparar modelo" panel. Click it and wait for the download to complete (this may take several minutes depending on your internet connection).
Alternatively, use Demo Mode ("Probar sin modelo") to test the interface without downloading the model. Results in demo mode are clearly marked as simulated.
The web interface accepts JPG, PNG and WEBP up to 10 MB. PDF is not supported by the interface pipeline.
curl -X POST "http://localhost:8000/api/ocr" \
-F "file=@document.jpg" \
-F "mode=markdown"
API Documentation: http://localhost:8000/docs
Detailed documentation is available in the /docs folder:
Interactive API documentation is available at http://localhost:8000/docs when the server is running.
deepseek-ocr/
โโโ ๐ Configuration Files
โ โโโ docker-compose.yml # Docker orchestration
โ โโโ .env.example # Environment template
โ โโโ .gitignore # Git ignore rules
โ
โโโ ๐ Documentation
โ โโโ README.md # Main documentation
โ โโโ LICENSE # MIT License (Dev/Test)
โ โโโ CONTRIBUTING.md # Contribution guidelines
โ โโโ CODE_OF_CONDUCT.md # Code of conduct
โ โโโ SECURITY.md # Security policy
โ โโโ docs/ # Additional documentation
โ
โโโ ๐ Backend (FastAPI)
โ โโโ main.py # API endpoints (thin shells)
โ โโโ model_lifecycle.py # Model loading state machine + DeepSeek adapter
โ โโโ ocr_handler.py # OCR request pipeline
โ โโโ config.py # Configuration
โ โโโ Dockerfile # Container image
โ โโโ requirements.txt # Python dependencies
โ โโโ tests/ # Unit tests (pytest, no GPU needed)
โ
โโโ ๐ Frontend (HTML/JS/CSS)
โ โโโ index.html # UI structure
โ โโโ app.js # Application logic
โ โโโ styles.css # Styling
โ โโโ nginx.conf # Web server config
โ โโโ Dockerfile # Container image
โ
โโโ ๐พ Data Directories
โ โโโ uploads/ # Uploaded images
โ โโโ outputs/ # OCR results
โ
โโโ ๐งช Testing
โโโ test_api.py # API test script
Edit docker-compose.yml or create a .env file to customize:
environment:
- CUDA_VISIBLE_DEVICES=0 # GPU to use
- MODEL_NAME=deepseek-ai/DeepSeek-OCR-2
- BASE_SIZE=1024 # Base resolution
- IMAGE_SIZE=768 # Crop size (OCR-2 default)
curl -X POST "http://localhost:8000/api/ocr" \
-F "file=@image.jpg" \
-F "mode=markdown"
| Mode | Interface label (Spanish) | Recommended Use |
|---|---|---|
markdown | Documento con formato | Documents (default) |
free_ocr | Texto sin formato | General text |
grounding | Texto con posiciones | Detailed analysis |
parse_figure | Figuras y grรกficos | Charts, tables |
detailed | Descripciรณn de la imagen | Visual analysis |
The web interface shows these as "Opciones avanzadas" (collapsed by default); the API accepts the raw values.
{
"text": "# Document Title\n\nExtracted content...",
"mode": "markdown",
"processing_time": 2.5,
"image_size": [1024, 768],
"tokens": 2257
}
# Document
"<image>\n<|grounding|>Convert the document to markdown."
# General image
"<image>\n<|grounding|>OCR this image."
# No format
"<image>\nFree OCR."
# Figures
"<image>\nParse the figure."
# Detailed description
"<image>\nDescribe this image in detail."
# Start services
docker-compose up -d
# View logs
docker-compose logs -f
# Stop services
docker-compose down
# Restart
docker-compose restart
# Rebuild images
docker-compose build --no-cache
curl http://localhost:8000/health
docker-compose logs -f deepseek-ocr-api
Benchmark results with 3503ร1668 pixels image on NVIDIA A100 40GB:
| Mode | Time | Quality | Structure |
|---|---|---|---|
| Free OCR | ~24s | โญโญโญ | Basic |
| Markdown | ~39s | โญโญโญ | Complete |
| Grounding | ~58s | โญโญ | + Coords |
| Detailed | ~9s | N/A | Description |
Hardware: NVIDIA A100 40GB
# Verify NVIDIA runtime
docker run --rm --gpus all nvidia/cuda:11.8.0-base-ubuntu22.04 nvidia-smi
docker-compose logs -f deepseek-ocr-apiReduce resolution in .env:
BASE_SIZE=640
IMAGE_SIZE=512
Change ports in docker-compose.yml:
ports:
- "3001:80" # Frontend (change 3000 to 3001)
- "8001:8000" # Backend (change 8000 to 8001)
For more help, check the documentation or open an issue.
The stack runs on hosts without an NVIDIA GPU (model loads on CPU, slower):
docker compose up -d
On Apple Silicon Macs, the backend forces platform: linux/amd64 because PyTorch cu118 wheels are x86_64-only. Note that loading the 3B-parameter model on CPU needs ~8GB RAM โ increase Docker Desktop's VM memory accordingly. Use Demo Mode in the web UI to validate the interface without loading the model.
MIT License - Development and Testing Only
This software is licensed under the MIT License with specific restrictions for development and testing purposes only. It is NOT intended for production use.
โ ๏ธ Production Use Warning: If you choose to use this software in production, you do so entirely at your own risk and responsibility. The authors provide no guarantees, support, or liability for production deployments.
See the LICENSE file for full terms and conditions.
requirements.txtContributions are welcome! Please read our Contributing Guidelines before submitting PRs.
git checkout -b feature/amazing-feature)git commit -m 'feat: add amazing feature')git push origin feature/amazing-feature)Please follow our Code of Conduct in all interactions.
โ ๏ธ This project is for development and testing only.
For security concerns, please review our Security Policy.
Key Security Notes:
security labelVersion: 2.0.0
Status: Active Development
Last Updated: September 2026
Model: DeepSeek-OCR-2 (deepseek-ai)
Purpose: Development and Testing Only
If you find this project helpful, please consider:
Made with โค๏ธ for the AI community