zademy/deepSeek-ocr-docker-compose

Python

21

12 commits

updated Oct 4, 2026

See the code

README

๐Ÿš€ DeepSeek OCR - AI-Powered Text Recognition

Complete OCR system using the DeepSeek-OCR-2 model (Released Jan 2026) with modern web interface and production-ready REST API.

License: MIT Docker Python FastAPI CUDA Code of Conduct

โš ๏ธ IMPORTANT: This project is for DEVELOPMENT and TESTING ONLY. Not intended for production use. See LICENSE for details.

โœจ Features

  • ๐Ÿค– Latest AI Model - DeepSeek-OCR-2 optimized for text recognition
  • ๐ŸŒ Interfaz clara - Flujo de tres pasos pensado para personas no tรฉcnicas
  • ๐Ÿ“Š Progress Tracking - Real-time model download progress bar
  • ๐ŸŽฎ Demo Mode - Test the interface without downloading the model
  • ๐Ÿ”Œ Complete REST API - Easy integration with FastAPI
  • ๐Ÿณ Docker Compose - Deploy in minutes with one command
  • โšก GPU Accelerated - NVIDIA CUDA support for maximum speed
  • โ™ฟ Accesible - Keyboard navigation, WCAG AA contrast, responsive from 320px
  • ๐Ÿ”“ 100% Open Source - MIT License for development/testing

๐Ÿ“š Table of Contents


๐Ÿ“ Requirements

  • Docker 20.10+ and Docker Compose 2.0+
  • NVIDIA GPU with CUDA 11.8+ (for GPU acceleration, optional โ€” see CPU-only mode)
  • At least 8GB VRAM (recommended for optimal performance)
  • 10GB disk space (for model cache)
  • Windows 10/11, Linux, or macOS (with Docker Desktop)

๐Ÿš€ Quick Start

1. Clone the Repository

git clone https://github.com/zademy/deepSeek-ocr-docker-compose
cd deepSeek-ocr-docker-compose

2. Configure Environment

cp .env.example .env
# Edit .env if needed (optional)

3. Start Services

# Standard (no GPU reservation โ€” works on any Docker host)
docker compose up -d

# With NVIDIA GPU acceleration
docker compose -f docker-compose.yml -f docker-compose.gpu.yml up -d

4. Access the Application

5. Download Model (First Time)

When you first access the web interface, you'll see a "Preparar modelo" panel. Click it and wait for the download to complete (this may take several minutes depending on your internet connection).

Alternatively, use Demo Mode ("Probar sin modelo") to test the interface without downloading the model. Results in demo mode are clearly marked as simulated.

Web interface formats

The web interface accepts JPG, PNG and WEBP up to 10 MB. PDF is not supported by the interface pipeline.

API Usage Example

curl -X POST "http://localhost:8000/api/ocr" \
  -F "file=@document.jpg" \
  -F "mode=markdown"

API Documentation: http://localhost:8000/docs

๐Ÿ“š Documentation

Detailed documentation is available in the /docs folder:

API Documentation

Interactive API documentation is available at http://localhost:8000/docs when the server is running.


๐Ÿงฐ Architecture

deepseek-ocr/
โ”œโ”€โ”€ ๐Ÿ“„ Configuration Files
โ”‚   โ”œโ”€โ”€ docker-compose.yml     # Docker orchestration
โ”‚   โ”œโ”€โ”€ .env.example           # Environment template
โ”‚   โ””โ”€โ”€ .gitignore             # Git ignore rules
โ”‚
โ”œโ”€โ”€ ๐Ÿ“– Documentation
โ”‚   โ”œโ”€โ”€ README.md              # Main documentation
โ”‚   โ”œโ”€โ”€ LICENSE                # MIT License (Dev/Test)
โ”‚   โ”œโ”€โ”€ CONTRIBUTING.md        # Contribution guidelines
โ”‚   โ”œโ”€โ”€ CODE_OF_CONDUCT.md     # Code of conduct
โ”‚   โ”œโ”€โ”€ SECURITY.md            # Security policy
โ”‚   โ””โ”€โ”€ docs/                  # Additional documentation
โ”‚
โ”œโ”€โ”€ ๐Ÿ Backend (FastAPI)
โ”‚   โ”œโ”€โ”€ main.py                # API endpoints (thin shells)
โ”‚   โ”œโ”€โ”€ model_lifecycle.py     # Model loading state machine + DeepSeek adapter
โ”‚   โ”œโ”€โ”€ ocr_handler.py         # OCR request pipeline
โ”‚   โ”œโ”€โ”€ config.py              # Configuration
โ”‚   โ”œโ”€โ”€ Dockerfile             # Container image
โ”‚   โ”œโ”€โ”€ requirements.txt       # Python dependencies
โ”‚   โ””โ”€โ”€ tests/                 # Unit tests (pytest, no GPU needed)
โ”‚
โ”œโ”€โ”€ ๐ŸŒ Frontend (HTML/JS/CSS)
โ”‚   โ”œโ”€โ”€ index.html             # UI structure
โ”‚   โ”œโ”€โ”€ app.js                 # Application logic
โ”‚   โ”œโ”€โ”€ styles.css             # Styling
โ”‚   โ”œโ”€โ”€ nginx.conf             # Web server config
โ”‚   โ””โ”€โ”€ Dockerfile             # Container image
โ”‚
โ”œโ”€โ”€ ๐Ÿ’พ Data Directories
โ”‚   โ”œโ”€โ”€ uploads/               # Uploaded images
โ”‚   โ””โ”€โ”€ outputs/               # OCR results
โ”‚
โ””โ”€โ”€ ๐Ÿงช Testing
    โ””โ”€โ”€ test_api.py            # API test script

๐Ÿ”ง Configuration

Environment Variables

Edit docker-compose.yml or create a .env file to customize:

environment:
  - CUDA_VISIBLE_DEVICES=0              # GPU to use
  - MODEL_NAME=deepseek-ai/DeepSeek-OCR-2
  - BASE_SIZE=1024                       # Base resolution
  - IMAGE_SIZE=768                       # Crop size (OCR-2 default)

๐Ÿ“– API Usage

Image OCR Endpoint

curl -X POST "http://localhost:8000/api/ocr" \
  -F "file=@image.jpg" \
  -F "mode=markdown"

Available Modes

ModeInterface label (Spanish)Recommended Use
markdownDocumento con formatoDocuments (default)
free_ocrTexto sin formatoGeneral text
groundingTexto con posicionesDetailed analysis
parse_figureFiguras y grรกficosCharts, tables
detailedDescripciรณn de la imagenVisual analysis

The web interface shows these as "Opciones avanzadas" (collapsed by default); the API accepts the raw values.

Response Example

{
  "text": "# Document Title\n\nExtracted content...",
  "mode": "markdown",
  "processing_time": 2.5,
  "image_size": [1024, 768],
  "tokens": 2257
}

๐ŸŽฏ Prompt Examples

# Document
"<image>\n<|grounding|>Convert the document to markdown."

# General image
"<image>\n<|grounding|>OCR this image."

# No format
"<image>\nFree OCR."

# Figures
"<image>\nParse the figure."

# Detailed description
"<image>\nDescribe this image in detail."

๐Ÿณ Docker Commands

# Start services
docker-compose up -d

# View logs
docker-compose logs -f

# Stop services
docker-compose down

# Restart
docker-compose restart

# Rebuild images
docker-compose build --no-cache

๐Ÿ” Monitoring

Health Check

curl http://localhost:8000/health

API Logs

docker-compose logs -f deepseek-ocr-api

๐Ÿ“Š Performance

Benchmark results with 3503ร—1668 pixels image on NVIDIA A100 40GB:

ModeTimeQualityStructure
Free OCR~24sโญโญโญBasic
Markdown~39sโญโญโญComplete
Grounding~58sโญโญ+ Coords
Detailed~9sN/ADescription

Hardware: NVIDIA A100 40GB

๐Ÿ› ๏ธ Supported Resolutions (DeepSeek-OCR-2)

  • Dynamic resolution (default): (0-6)ร—768ร—768 + 1ร—1024ร—1024 โ€” (0-6)ร—144 + 256 visual tokens

๐Ÿ› Troubleshooting

GPU Not Detected

# Verify NVIDIA runtime
docker run --rm --gpus all nvidia/cuda:11.8.0-base-ubuntu22.04 nvidia-smi

Model Not Downloading

  • Check internet connection
  • Verify disk space (need ~7GB free)
  • Use the download button in the web interface
  • Check logs: docker-compose logs -f deepseek-ocr-api

Out of Memory

Reduce resolution in .env:

BASE_SIZE=640
IMAGE_SIZE=512

Port Already in Use

Change ports in docker-compose.yml:

ports:
  - "3001:80"  # Frontend (change 3000 to 3001)
  - "8001:8000"  # Backend (change 8000 to 8001)

For more help, check the documentation or open an issue.

๐Ÿ’ป CPU-Only / Apple Silicon

The stack runs on hosts without an NVIDIA GPU (model loads on CPU, slower):

docker compose up -d

On Apple Silicon Macs, the backend forces platform: linux/amd64 because PyTorch cu118 wheels are x86_64-only. Note that loading the 3B-parameter model on CPU needs ~8GB RAM โ€” increase Docker Desktop's VM memory accordingly. Use Demo Mode in the web UI to validate the interface without loading the model.

๐Ÿ“œ Resources


๐Ÿ“ License

MIT License - Development and Testing Only

This software is licensed under the MIT License with specific restrictions for development and testing purposes only. It is NOT intended for production use.

โš ๏ธ Production Use Warning: If you choose to use this software in production, you do so entirely at your own risk and responsibility. The authors provide no guarantees, support, or liability for production deployments.

See the LICENSE file for full terms and conditions.

Third-Party Components

  • DeepSeek-OCR-2 Model: Subject to its own license terms
  • Other dependencies: Check individual package licenses in requirements.txt

๐Ÿค Contributing

Contributions are welcome! Please read our Contributing Guidelines before submitting PRs.

How to Contribute

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'feat: add amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

Please follow our Code of Conduct in all interactions.


๐Ÿ”’ Security

โš ๏ธ This project is for development and testing only.

For security concerns, please review our Security Policy.

Key Security Notes:

  • No authentication implemented
  • Not hardened for production use
  • Use at your own risk in production environments
  • Report vulnerabilities via GitHub issues with security label

๐Ÿš€ Getting Help

  • Documentation: Check the docs folder
  • Issues: Open an issue on GitHub
  • Discussions: Use GitHub Discussions for questions
  • API Docs: Visit http://localhost:8000/docs when running

๐Ÿ“Œ Project Status

Version: 2.0.0
Status: Active Development
Last Updated: September 2026 Model: DeepSeek-OCR-2 (deepseek-ai)
Purpose: Development and Testing Only


โญ Show Your Support

If you find this project helpful, please consider:

  • Giving it a โญ on GitHub
  • Sharing it with others
  • Contributing improvements
  • Reporting bugs and suggestions

Made with โค๏ธ for the AI community

zademy/deepSeek-ocr-docker-compose

Python

21

12 commits

updated Oct 4, 2026

See the code

README

๐Ÿš€ DeepSeek OCR - AI-Powered Text Recognition

Complete OCR system using the DeepSeek-OCR-2 model (Released Jan 2026) with modern web interface and production-ready REST API.

License: MIT Docker Python FastAPI CUDA Code of Conduct

โš ๏ธ IMPORTANT: This project is for DEVELOPMENT and TESTING ONLY. Not intended for production use. See LICENSE for details.

โœจ Features

  • ๐Ÿค– Latest AI Model - DeepSeek-OCR-2 optimized for text recognition
  • ๐ŸŒ Interfaz clara - Flujo de tres pasos pensado para personas no tรฉcnicas
  • ๐Ÿ“Š Progress Tracking - Real-time model download progress bar
  • ๐ŸŽฎ Demo Mode - Test the interface without downloading the model
  • ๐Ÿ”Œ Complete REST API - Easy integration with FastAPI
  • ๐Ÿณ Docker Compose - Deploy in minutes with one command
  • โšก GPU Accelerated - NVIDIA CUDA support for maximum speed
  • โ™ฟ Accesible - Keyboard navigation, WCAG AA contrast, responsive from 320px
  • ๐Ÿ”“ 100% Open Source - MIT License for development/testing

๐Ÿ“š Table of Contents


๐Ÿ“ Requirements

  • Docker 20.10+ and Docker Compose 2.0+
  • NVIDIA GPU with CUDA 11.8+ (for GPU acceleration, optional โ€” see CPU-only mode)
  • At least 8GB VRAM (recommended for optimal performance)
  • 10GB disk space (for model cache)
  • Windows 10/11, Linux, or macOS (with Docker Desktop)

๐Ÿš€ Quick Start

1. Clone the Repository

git clone https://github.com/zademy/deepSeek-ocr-docker-compose
cd deepSeek-ocr-docker-compose

2. Configure Environment

cp .env.example .env
# Edit .env if needed (optional)

3. Start Services

# Standard (no GPU reservation โ€” works on any Docker host)
docker compose up -d

# With NVIDIA GPU acceleration
docker compose -f docker-compose.yml -f docker-compose.gpu.yml up -d

4. Access the Application

5. Download Model (First Time)

When you first access the web interface, you'll see a "Preparar modelo" panel. Click it and wait for the download to complete (this may take several minutes depending on your internet connection).

Alternatively, use Demo Mode ("Probar sin modelo") to test the interface without downloading the model. Results in demo mode are clearly marked as simulated.

Web interface formats

The web interface accepts JPG, PNG and WEBP up to 10 MB. PDF is not supported by the interface pipeline.

API Usage Example

curl -X POST "http://localhost:8000/api/ocr" \
  -F "file=@document.jpg" \
  -F "mode=markdown"

API Documentation: http://localhost:8000/docs

๐Ÿ“š Documentation

Detailed documentation is available in the /docs folder:

API Documentation

Interactive API documentation is available at http://localhost:8000/docs when the server is running.


๐Ÿงฐ Architecture

deepseek-ocr/
โ”œโ”€โ”€ ๐Ÿ“„ Configuration Files
โ”‚   โ”œโ”€โ”€ docker-compose.yml     # Docker orchestration
โ”‚   โ”œโ”€โ”€ .env.example           # Environment template
โ”‚   โ””โ”€โ”€ .gitignore             # Git ignore rules
โ”‚
โ”œโ”€โ”€ ๐Ÿ“– Documentation
โ”‚   โ”œโ”€โ”€ README.md              # Main documentation
โ”‚   โ”œโ”€โ”€ LICENSE                # MIT License (Dev/Test)
โ”‚   โ”œโ”€โ”€ CONTRIBUTING.md        # Contribution guidelines
โ”‚   โ”œโ”€โ”€ CODE_OF_CONDUCT.md     # Code of conduct
โ”‚   โ”œโ”€โ”€ SECURITY.md            # Security policy
โ”‚   โ””โ”€โ”€ docs/                  # Additional documentation
โ”‚
โ”œโ”€โ”€ ๐Ÿ Backend (FastAPI)
โ”‚   โ”œโ”€โ”€ main.py                # API endpoints (thin shells)
โ”‚   โ”œโ”€โ”€ model_lifecycle.py     # Model loading state machine + DeepSeek adapter
โ”‚   โ”œโ”€โ”€ ocr_handler.py         # OCR request pipeline
โ”‚   โ”œโ”€โ”€ config.py              # Configuration
โ”‚   โ”œโ”€โ”€ Dockerfile             # Container image
โ”‚   โ”œโ”€โ”€ requirements.txt       # Python dependencies
โ”‚   โ””โ”€โ”€ tests/                 # Unit tests (pytest, no GPU needed)
โ”‚
โ”œโ”€โ”€ ๐ŸŒ Frontend (HTML/JS/CSS)
โ”‚   โ”œโ”€โ”€ index.html             # UI structure
โ”‚   โ”œโ”€โ”€ app.js                 # Application logic
โ”‚   โ”œโ”€โ”€ styles.css             # Styling
โ”‚   โ”œโ”€โ”€ nginx.conf             # Web server config
โ”‚   โ””โ”€โ”€ Dockerfile             # Container image
โ”‚
โ”œโ”€โ”€ ๐Ÿ’พ Data Directories
โ”‚   โ”œโ”€โ”€ uploads/               # Uploaded images
โ”‚   โ””โ”€โ”€ outputs/               # OCR results
โ”‚
โ””โ”€โ”€ ๐Ÿงช Testing
    โ””โ”€โ”€ test_api.py            # API test script

๐Ÿ”ง Configuration

Environment Variables

Edit docker-compose.yml or create a .env file to customize:

environment:
  - CUDA_VISIBLE_DEVICES=0              # GPU to use
  - MODEL_NAME=deepseek-ai/DeepSeek-OCR-2
  - BASE_SIZE=1024                       # Base resolution
  - IMAGE_SIZE=768                       # Crop size (OCR-2 default)

๐Ÿ“– API Usage

Image OCR Endpoint

curl -X POST "http://localhost:8000/api/ocr" \
  -F "file=@image.jpg" \
  -F "mode=markdown"

Available Modes

ModeInterface label (Spanish)Recommended Use
markdownDocumento con formatoDocuments (default)
free_ocrTexto sin formatoGeneral text
groundingTexto con posicionesDetailed analysis
parse_figureFiguras y grรกficosCharts, tables
detailedDescripciรณn de la imagenVisual analysis

The web interface shows these as "Opciones avanzadas" (collapsed by default); the API accepts the raw values.

Response Example

{
  "text": "# Document Title\n\nExtracted content...",
  "mode": "markdown",
  "processing_time": 2.5,
  "image_size": [1024, 768],
  "tokens": 2257
}

๐ŸŽฏ Prompt Examples

# Document
"<image>\n<|grounding|>Convert the document to markdown."

# General image
"<image>\n<|grounding|>OCR this image."

# No format
"<image>\nFree OCR."

# Figures
"<image>\nParse the figure."

# Detailed description
"<image>\nDescribe this image in detail."

๐Ÿณ Docker Commands

# Start services
docker-compose up -d

# View logs
docker-compose logs -f

# Stop services
docker-compose down

# Restart
docker-compose restart

# Rebuild images
docker-compose build --no-cache

๐Ÿ” Monitoring

Health Check

curl http://localhost:8000/health

API Logs

docker-compose logs -f deepseek-ocr-api

๐Ÿ“Š Performance

Benchmark results with 3503ร—1668 pixels image on NVIDIA A100 40GB:

ModeTimeQualityStructure
Free OCR~24sโญโญโญBasic
Markdown~39sโญโญโญComplete
Grounding~58sโญโญ+ Coords
Detailed~9sN/ADescription

Hardware: NVIDIA A100 40GB

๐Ÿ› ๏ธ Supported Resolutions (DeepSeek-OCR-2)

  • Dynamic resolution (default): (0-6)ร—768ร—768 + 1ร—1024ร—1024 โ€” (0-6)ร—144 + 256 visual tokens

๐Ÿ› Troubleshooting

GPU Not Detected

# Verify NVIDIA runtime
docker run --rm --gpus all nvidia/cuda:11.8.0-base-ubuntu22.04 nvidia-smi

Model Not Downloading

  • Check internet connection
  • Verify disk space (need ~7GB free)
  • Use the download button in the web interface
  • Check logs: docker-compose logs -f deepseek-ocr-api

Out of Memory

Reduce resolution in .env:

BASE_SIZE=640
IMAGE_SIZE=512

Port Already in Use

Change ports in docker-compose.yml:

ports:
  - "3001:80"  # Frontend (change 3000 to 3001)
  - "8001:8000"  # Backend (change 8000 to 8001)

For more help, check the documentation or open an issue.

๐Ÿ’ป CPU-Only / Apple Silicon

The stack runs on hosts without an NVIDIA GPU (model loads on CPU, slower):

docker compose up -d

On Apple Silicon Macs, the backend forces platform: linux/amd64 because PyTorch cu118 wheels are x86_64-only. Note that loading the 3B-parameter model on CPU needs ~8GB RAM โ€” increase Docker Desktop's VM memory accordingly. Use Demo Mode in the web UI to validate the interface without loading the model.

๐Ÿ“œ Resources


๐Ÿ“ License

MIT License - Development and Testing Only

This software is licensed under the MIT License with specific restrictions for development and testing purposes only. It is NOT intended for production use.

โš ๏ธ Production Use Warning: If you choose to use this software in production, you do so entirely at your own risk and responsibility. The authors provide no guarantees, support, or liability for production deployments.

See the LICENSE file for full terms and conditions.

Third-Party Components

  • DeepSeek-OCR-2 Model: Subject to its own license terms
  • Other dependencies: Check individual package licenses in requirements.txt

๐Ÿค Contributing

Contributions are welcome! Please read our Contributing Guidelines before submitting PRs.

How to Contribute

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'feat: add amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

Please follow our Code of Conduct in all interactions.


๐Ÿ”’ Security

โš ๏ธ This project is for development and testing only.

For security concerns, please review our Security Policy.

Key Security Notes:

  • No authentication implemented
  • Not hardened for production use
  • Use at your own risk in production environments
  • Report vulnerabilities via GitHub issues with security label

๐Ÿš€ Getting Help

  • Documentation: Check the docs folder
  • Issues: Open an issue on GitHub
  • Discussions: Use GitHub Discussions for questions
  • API Docs: Visit http://localhost:8000/docs when running

๐Ÿ“Œ Project Status

Version: 2.0.0
Status: Active Development
Last Updated: September 2026 Model: DeepSeek-OCR-2 (deepseek-ai)
Purpose: Development and Testing Only


โญ Show Your Support

If you find this project helpful, please consider:

  • Giving it a โญ on GitHub
  • Sharing it with others
  • Contributing improvements
  • Reporting bugs and suggestions

Made with โค๏ธ for the AI community