One stop shop for running AI/ML on AWS.
1,189
stars
666
commits
Python
primary language
Sep 10, 2026
updated
One stop shop for running AI/ML on AWS
Docs · Available Images · Tutorials
AWS Deep Learning Containers (DLCs) are pre-built Docker images for running AI/ML workloads on AWS. Each image is tested and patched for security vulnerabilities. For more details, visit our documentation.
0.5.19-gpu-py312-ec2 · SageMaker: 0.5.19-gpu-py312 · Qwen3.8, Ling-3.0, Spark2.5, Granite 4.2; beam search; DeepEP v2 MoE all-to-all.serve-llm-cuda-v1.0 · Initial release: OpenAI-compatible LLM serving with Ray Serve and vLLM 0.26.0 on PyTorch 2.11.0 / CUDA 13.0.2 / Python 3.13; ray[llm]'s build_openai_app runs vLLM behind Ray Serve — single GPU on EC2, and multi-node serving on EKS via KubeRay.train-ml-cuda-v1.1 · EFA 1.49.0 (up from 1.47.0).0.28.0-gpu-py312-ec2 · SageMaker: 0.28.0-gpu-py312 · Kimi-K3 stack-wide optimization (Decode Context Parallel, fused FlashKDA kernels, GEMM-RS sequence parallelism); DeepSeek V4 sparse MLA end-to-end with MTP and DSpark speculative decoding; new models Muse Glimmer, Ling 3.0 Flash, Dots3, Interns2mobius; tiered KV cache offloading to disk; runtime base moves to Ubuntu 24.04 and Transformers 5.15.0; max_num_batched_tokens default 8192 -> 16384.train-ml-cuda-v1.0 · Initial release: multi-node, multi-GPU distributed training with Ray Train/Tune/Data on PyTorch 2.13.0 / CUDA 13.0.2 / Python 3.13, with EFA 1.47.0, the AWS NCCL OFI plugin, GDRCopy, flash-attn, Transformer Engine, and DeepSpeed pre-installed; one image runs as either a Ray head or worker under KubeRay on any EKS cluster (including SageMaker HyperPod-EKS), or standalone on EC2.omni-cuda-v1.6 · SageMaker: omni-sagemaker-cuda-v1.6 · SageMaker SM_VLLM_* fix — JSON-array env vars (e.g. SM_VLLM_LORA_MODULES) now expand into multiple argv values so multi-value flags parse correctly; no framework bump (still vLLM-Omni 0.26.0).server-cuda-v2.4 · SageMaker: server-sagemaker-cuda-v2.4 · vLLM 0.27.1 (up from 0.27.0); Muse Glimmer model support; SageMaker SM_VLLM_* fix — JSON-array env vars (e.g. SM_VLLM_LORA_MODULES) now expand into multiple argv values so multi-value flags parse correctly.0.5.18-gpu-py312-ec2 · SageMaker: 0.5.18-gpu-py312 · Muse Glimmer and Intern-S2-Mobius, plus diffusion additions SANA-Video, LTX-2.5, Cosmos3 Edge, and LongCat-Image; overlapped checkpoint staging for faster startup (--startup-weight-load-mode overlap); FlashInfer MNNVL standalone allreduce on by default for DeepSeek-V3/V3.2/V4; upstream stack moves to torch 2.13.0 and triton 3.7.1.server-cpu-v1 · server-cuda-v1 · Graviton: llama-cpp-arm64:server-cpu-v1 · SageMaker: server-sagemaker-cpu-v1 · server-sagemaker-cuda-v1 · llama-cpp-arm64:server-sagemaker-cpu-v1 · Initial release: serve quantized GGUF models with the upstream llama-server OpenAI-compatible API on x86 CPU, NVIDIA GPU (CUDA 13.0.2), and AWS Graviton (ARM64); Python 3.12.2.20.0-cpu-py312-amzn2023-sagemaker · 2.20.0-gpu-py312-cu129-amzn2023-sagemaker · TensorFlow Serving 2.20.0 on Amazon Linux 2023 with Python 3.12 and CUDA 12.9.server-cuda-v1.3 · SageMaker: server-sagemaker-cuda-v1.3 · SGLang 0.5.17 (up from 0.5.14); Kimi-K3 (2.8T MoE, MXFP4) support; sgl-kernel 0.4.5, FlashInfer 0.6.15.post1, Mooncake 0.3.12.post1.server-cuda-v2.3 · SageMaker: server-sagemaker-cuda-v2.3 · vLLM 0.27.0 (up from 0.26.0); Kimi K3 (native support + kernels, Rust/Python frontends); FlashInfer 0.6.16.post3; NVIDIA B300 (SM103); new models K-EXAONE-2.0-750B-A37B, jina-embeddings-v5-text-nano, Qwen3.5; dynamic FP8 for Inkling; Baidu Unlimited-OCR smoke test.0.27.1-gpu-py312-ec2 · SageMaker: 0.27.1-gpu-py312 · Kimi K3, Qwen3.5 dense + MoE (EVS video token pruning), K-EXAONE-2.0-750B-A37B, VaultGemma, jina-embeddings-v5-text-nano.3.8.6-cu128-amzn2023 · SageMaker: 3.8.6-cu128-amzn2023-sagemaker · Initial release: speech transcription with word-level alignment (wav2vec2) and speaker diarization (pyannote) through an OpenAI-compatible API on CUDA 12.8 / Python 3.12; real-time and asynchronous SageMaker endpoints.serve-ml-cuda-v1.4 · serve-ml-cpu-v1.4 · SageMaker: serve-ml-sagemaker-cuda-v1.4 · serve-ml-sagemaker-cpu-v1.4 · Ray 2.57.0 (up from 2.56.1).This project is licensed under the Apache-2.0 License.
111 commits
79 commits
71 commits
71 commits
Python
71.1%
Shell
22.1%
Dockerfile
6.1%
One stop shop for running AI/ML on AWS.
1,189
stars
666
commits
Python
primary language
Sep 10, 2026
updated
One stop shop for running AI/ML on AWS
Docs · Available Images · Tutorials
AWS Deep Learning Containers (DLCs) are pre-built Docker images for running AI/ML workloads on AWS. Each image is tested and patched for security vulnerabilities. For more details, visit our documentation.
0.5.19-gpu-py312-ec2 · SageMaker: 0.5.19-gpu-py312 · Qwen3.8, Ling-3.0, Spark2.5, Granite 4.2; beam search; DeepEP v2 MoE all-to-all.serve-llm-cuda-v1.0 · Initial release: OpenAI-compatible LLM serving with Ray Serve and vLLM 0.26.0 on PyTorch 2.11.0 / CUDA 13.0.2 / Python 3.13; ray[llm]'s build_openai_app runs vLLM behind Ray Serve — single GPU on EC2, and multi-node serving on EKS via KubeRay.train-ml-cuda-v1.1 · EFA 1.49.0 (up from 1.47.0).0.28.0-gpu-py312-ec2 · SageMaker: 0.28.0-gpu-py312 · Kimi-K3 stack-wide optimization (Decode Context Parallel, fused FlashKDA kernels, GEMM-RS sequence parallelism); DeepSeek V4 sparse MLA end-to-end with MTP and DSpark speculative decoding; new models Muse Glimmer, Ling 3.0 Flash, Dots3, Interns2mobius; tiered KV cache offloading to disk; runtime base moves to Ubuntu 24.04 and Transformers 5.15.0; max_num_batched_tokens default 8192 -> 16384.train-ml-cuda-v1.0 · Initial release: multi-node, multi-GPU distributed training with Ray Train/Tune/Data on PyTorch 2.13.0 / CUDA 13.0.2 / Python 3.13, with EFA 1.47.0, the AWS NCCL OFI plugin, GDRCopy, flash-attn, Transformer Engine, and DeepSpeed pre-installed; one image runs as either a Ray head or worker under KubeRay on any EKS cluster (including SageMaker HyperPod-EKS), or standalone on EC2.omni-cuda-v1.6 · SageMaker: omni-sagemaker-cuda-v1.6 · SageMaker SM_VLLM_* fix — JSON-array env vars (e.g. SM_VLLM_LORA_MODULES) now expand into multiple argv values so multi-value flags parse correctly; no framework bump (still vLLM-Omni 0.26.0).server-cuda-v2.4 · SageMaker: server-sagemaker-cuda-v2.4 · vLLM 0.27.1 (up from 0.27.0); Muse Glimmer model support; SageMaker SM_VLLM_* fix — JSON-array env vars (e.g. SM_VLLM_LORA_MODULES) now expand into multiple argv values so multi-value flags parse correctly.0.5.18-gpu-py312-ec2 · SageMaker: 0.5.18-gpu-py312 · Muse Glimmer and Intern-S2-Mobius, plus diffusion additions SANA-Video, LTX-2.5, Cosmos3 Edge, and LongCat-Image; overlapped checkpoint staging for faster startup (--startup-weight-load-mode overlap); FlashInfer MNNVL standalone allreduce on by default for DeepSeek-V3/V3.2/V4; upstream stack moves to torch 2.13.0 and triton 3.7.1.server-cpu-v1 · server-cuda-v1 · Graviton: llama-cpp-arm64:server-cpu-v1 · SageMaker: server-sagemaker-cpu-v1 · server-sagemaker-cuda-v1 · llama-cpp-arm64:server-sagemaker-cpu-v1 · Initial release: serve quantized GGUF models with the upstream llama-server OpenAI-compatible API on x86 CPU, NVIDIA GPU (CUDA 13.0.2), and AWS Graviton (ARM64); Python 3.12.2.20.0-cpu-py312-amzn2023-sagemaker · 2.20.0-gpu-py312-cu129-amzn2023-sagemaker · TensorFlow Serving 2.20.0 on Amazon Linux 2023 with Python 3.12 and CUDA 12.9.server-cuda-v1.3 · SageMaker: server-sagemaker-cuda-v1.3 · SGLang 0.5.17 (up from 0.5.14); Kimi-K3 (2.8T MoE, MXFP4) support; sgl-kernel 0.4.5, FlashInfer 0.6.15.post1, Mooncake 0.3.12.post1.server-cuda-v2.3 · SageMaker: server-sagemaker-cuda-v2.3 · vLLM 0.27.0 (up from 0.26.0); Kimi K3 (native support + kernels, Rust/Python frontends); FlashInfer 0.6.16.post3; NVIDIA B300 (SM103); new models K-EXAONE-2.0-750B-A37B, jina-embeddings-v5-text-nano, Qwen3.5; dynamic FP8 for Inkling; Baidu Unlimited-OCR smoke test.0.27.1-gpu-py312-ec2 · SageMaker: 0.27.1-gpu-py312 · Kimi K3, Qwen3.5 dense + MoE (EVS video token pruning), K-EXAONE-2.0-750B-A37B, VaultGemma, jina-embeddings-v5-text-nano.3.8.6-cu128-amzn2023 · SageMaker: 3.8.6-cu128-amzn2023-sagemaker · Initial release: speech transcription with word-level alignment (wav2vec2) and speaker diarization (pyannote) through an OpenAI-compatible API on CUDA 12.8 / Python 3.12; real-time and asynchronous SageMaker endpoints.serve-ml-cuda-v1.4 · serve-ml-cpu-v1.4 · SageMaker: serve-ml-sagemaker-cuda-v1.4 · serve-ml-sagemaker-cpu-v1.4 · Ray 2.57.0 (up from 2.56.1).This project is licensed under the Apache-2.0 License.
111 commits
79 commits
71 commits
71 commits
Python
71.1%
Shell
22.1%
Dockerfile
6.1%