aws-neuron/aws-neuron-eks-samples

Python

25

182 commits

updated Sep 8, 2026

See the code

README

AWS Neuron EKS Samples

This repository contains samples for Amazon Elastic Kubernetes Service (EKS) and AWS Neuron, the software development kit (SDK) that enables machine learning (ML) inference and training workloads on the AWS ML accelerator chips Inferentia and Trainium.

The samples in this repository demonstrate the types of patterns that can be used to deliver inference and distributed training on EKS using Inferentia and Trainium. The samples can be used as-is, or easily modified to support additional models and use cases.

Samples are organized by use case below:

Training

LinkDescriptionInstance Type
BERT pretrainingEnd-end workflow for creating an EKS cluster with 2 trn1.32xl nodes and running BERT phase1 pretraining (64-worker DataParallel)Trn1
MLP trainingIntroductory workflow for creating an EKS cluster with 1 node and running a simple MLP training jobTrn1
Llama 3.1 8B finetuning with Ray+PTLEnd-end workflow for creating a Ray cluster with 2 trn1.32xlarge nodes on EKS and running Llama 3.1 8B finetuningTrn1

Inference

LinkDescriptionInstance Type
SD inferenceSD Inference workflow for creating an inference endpoint forwarded by ALB LoadBalancer powered by Karpenter's NodePoolInf2
Flux inferenceFLUX.1 dev Inference workflow for creating an inference endpoint forwarded by ALB LoadBalancer powered by Karpenter's NodePool and S3 mountpointsTrn1/Inf2
Optimal TP/DP for LLM servingDemonstrates optimal tensor parallelism configuration for LLM serving with vLLM, comparing TP1, TP2, and TP4 performance on Qwen modelsTrn2
Speculative decodingAccelerate LLM inference using speculative decoding with vLLM, comparing baseline vs draft model performance with Neuron DRA and S3 persistenceTrn2
Disaggregated inferencePrefill/decode disaggregated serving with vLLM, using DRA for Neuron + EFA allocation and NIXL/LIBFABRIC KV transfer; dynamic xPyD scaling routed by vLLM production-stack or AIBrixTrn2/Trn3

Getting Help

If you encounter issues with any of the samples in this repository, please open an issue via the GitHub Issues feature.

Contributing

Please refer to the CONTRIBUTING document for details on contributing additional samples to this repository.

Release Notes

Please refer to the Change Log.

aws-neuron/aws-neuron-eks-samples

Python

25

182 commits

updated Sep 8, 2026

See the code

README

AWS Neuron EKS Samples

This repository contains samples for Amazon Elastic Kubernetes Service (EKS) and AWS Neuron, the software development kit (SDK) that enables machine learning (ML) inference and training workloads on the AWS ML accelerator chips Inferentia and Trainium.

The samples in this repository demonstrate the types of patterns that can be used to deliver inference and distributed training on EKS using Inferentia and Trainium. The samples can be used as-is, or easily modified to support additional models and use cases.

Samples are organized by use case below:

Training

LinkDescriptionInstance Type
BERT pretrainingEnd-end workflow for creating an EKS cluster with 2 trn1.32xl nodes and running BERT phase1 pretraining (64-worker DataParallel)Trn1
MLP trainingIntroductory workflow for creating an EKS cluster with 1 node and running a simple MLP training jobTrn1
Llama 3.1 8B finetuning with Ray+PTLEnd-end workflow for creating a Ray cluster with 2 trn1.32xlarge nodes on EKS and running Llama 3.1 8B finetuningTrn1

Inference

LinkDescriptionInstance Type
SD inferenceSD Inference workflow for creating an inference endpoint forwarded by ALB LoadBalancer powered by Karpenter's NodePoolInf2
Flux inferenceFLUX.1 dev Inference workflow for creating an inference endpoint forwarded by ALB LoadBalancer powered by Karpenter's NodePool and S3 mountpointsTrn1/Inf2
Optimal TP/DP for LLM servingDemonstrates optimal tensor parallelism configuration for LLM serving with vLLM, comparing TP1, TP2, and TP4 performance on Qwen modelsTrn2
Speculative decodingAccelerate LLM inference using speculative decoding with vLLM, comparing baseline vs draft model performance with Neuron DRA and S3 persistenceTrn2
Disaggregated inferencePrefill/decode disaggregated serving with vLLM, using DRA for Neuron + EFA allocation and NIXL/LIBFABRIC KV transfer; dynamic xPyD scaling routed by vLLM production-stack or AIBrixTrn2/Trn3

Getting Help

If you encounter issues with any of the samples in this repository, please open an issue via the GitHub Issues feature.

Contributing

Please refer to the CONTRIBUTING document for details on contributing additional samples to this repository.

Release Notes

Please refer to the Change Log.

Languages

Python

95.5%

Shell

2.7%