SalesforceAIResearch/CoDA

Salesforce AI Research's open diffusion language model

Python

67

3 commits

updated Jun 2, 2026

See the code

README

CoDA: Coding LM via Diffusion Adaptation

CoDA Logo

End-to-end diffusion language modeling across TPU pre-training, GPU fine-tuning, evaluation, and serving.

Paper Model


Table of Contents

Overview 🎯

CoDA is Salesforce AI Research's open diffusion language model. This repo contains a unified training pipeline from pre-training to post-training, evaluation harnesses, and a simple Fast-API based serving backend.

Note: This repository is provided for research purposes only. Data release is subjected to internal regulations.

Repository Map πŸ—ΊοΈ

DirectoryPurpose
CoDALanguageModel/Huggingface model class
post-train/Supervised fine-tuning (SFT) pipeline
evaluation/Evaluation framework
pre-train/TPU-based pre-training pipeline
serving/Serving stack
run_sft.shLauncher coupling pre-training checkpoints with the post-training diffusion trainer.
save_hf_model.pyUtil function to convert checkpoint in Huggingface model class

Training Quickstart πŸš€

To avoid dependency conflicts, we recommend maintaining isolated environments per subsystem and activate the corresponding environment before executing scripts in each subdirectory.

Pre-training on TPU (pre-train/)

  1. Populate TPU metadata in pre-train/env.example and copy to pre-train/.env.
  2. Run pre-train/setup_tpu.sh to provision dependencies and sync the repository to the TPU pod.
  3. Launch pre-training with the provided recipes (e.g., pre-train/recipes/midtrain_v4_512.sh) to produce CoDA checkpoints (GCS or local storage).

Supervised Diffusion Fine-tuning (post-train/)

  1. Install prerequisites following post-train/LLaMA-Factory/README.md.
  2. Configure dataset metadata in post-train/LLaMA-Factory/data/dataset_info.json and diffusion arguments in post-train/LLaMA-Factory/examples/train_full/*.yaml.
  3. Execute ./run_sft.sh to fine-tune CoDA checkpoints with discrete denoising objectives.

Evaluation Pipelines (evaluation/)

  1. Choose a benchmark script such as evaluation/lm_eval/eval_mbpp_humaneval.sh.
  2. Update MODEL_DIR and diffusion parameters (diffusion_steps, temperature, top_p) to match the target checkpoint.
  3. Run the script to gather metrics; logs are stored locally for aggregation and reporting.

Benchmark πŸ“Š

Comparison of code-generation performance across standard and plus-enhanced benchmarks. Evalplus is computed as the mean pass@1 on enhanced variants. Bold marks results where CoDA produces the strongest diffusion-model performance.

ModelHumaneval InstructHumaneval PlusMBPP InstructMBPP PlusEvalplus
CoDA-Base29.323.835.246.034.9
CoDA-Instruct54.347.647.263.255.4
Dream-Base56.750.068.757.453.7
Dream-7B-Instruct57.953.768.356.154.9
LLaDA-8B-Instruct35.431.731.528.630.2
Qwen3-1.7B66.561.646.265.963.8
Qwen2.5-Coder-1.5B43.936.669.258.647.6
Qwen2.5-Coder-1.5B-Instruct70.766.569.259.462.3
Gemma-3-1B-it39.635.439.463.549.5
LLaMA-3.2-1B-Instruct35.431.124.453.742.4

Deployment Guide πŸ› οΈ

python3 -m venv .venv
source .venv/bin/activate

2 Install dependencies

pip install -r serving/requirements.txt

3 Export your Hugging Face token

export HF_TOKEN="hf_..."

4 Serve the model on cuda

bash serving/fast-api/start_server.sh

The server will listen on http://localhost:8000.

5 Interact with the served model

python serving/fast-api/chat_cli.py --base-url http://localhost:8000 --model Salesforce/CoDA-v0-Instruct

Optional flags:

  • --stream to stream tokens as they are generated
  • --show-meta to display latency and token usage

Generation hyperparameters (env vars)

You can customize generation with these environment variables (defaults in parentheses):

  • MAX_TOKENS (256)
  • TEMPERATURE (0.0)
  • TOP_P (unset)
  • TOP_K (unset)
  • STEPS (128)
  • ALG ("entropy")
  • ALG_TEMP (0.1)
  • BLOCK_LENGTH (32)

Example:

export MAX_TOKENS=512
export TEMPERATURE=0.7
export TOP_P=0.9
export STEPS=128
export ALG=entropy
export ALG_TEMP=0.1
export BLOCK_LENGTH=32
bash serving/fast-api/start_server.sh

Citation πŸ“š

@misc{coda2025,
  title={CoDA: Coding LM via Diffusion Adaptation},
  author={Chen, Haolin and Wang, Shiyu and Qin, Can and Pang, Bo and Liu, Zuxin and Qiu, Jielin and Zhang, Jianguo and Zhou, Yingbo and Chen, Zeyuan and Xu, Ran and Heinecke, Shelby and Savarese, Silvio and Xiong, Caiming and Wang, Huan and Yao, Weiran},
  year={2025},
  publisher={Salesforce AI Research}
}
coding
diffusion-models
llm
post-training
pre-training
sft

SalesforceAIResearch/CoDA

Salesforce AI Research's open diffusion language model

Python

67

3 commits

updated Jun 2, 2026

See the code

README

CoDA: Coding LM via Diffusion Adaptation

CoDA Logo

End-to-end diffusion language modeling across TPU pre-training, GPU fine-tuning, evaluation, and serving.

Paper Model


Table of Contents

Overview 🎯

CoDA is Salesforce AI Research's open diffusion language model. This repo contains a unified training pipeline from pre-training to post-training, evaluation harnesses, and a simple Fast-API based serving backend.

Note: This repository is provided for research purposes only. Data release is subjected to internal regulations.

Repository Map πŸ—ΊοΈ

DirectoryPurpose
CoDALanguageModel/Huggingface model class
post-train/Supervised fine-tuning (SFT) pipeline
evaluation/Evaluation framework
pre-train/TPU-based pre-training pipeline
serving/Serving stack
run_sft.shLauncher coupling pre-training checkpoints with the post-training diffusion trainer.
save_hf_model.pyUtil function to convert checkpoint in Huggingface model class

Training Quickstart πŸš€

To avoid dependency conflicts, we recommend maintaining isolated environments per subsystem and activate the corresponding environment before executing scripts in each subdirectory.

Pre-training on TPU (pre-train/)

  1. Populate TPU metadata in pre-train/env.example and copy to pre-train/.env.
  2. Run pre-train/setup_tpu.sh to provision dependencies and sync the repository to the TPU pod.
  3. Launch pre-training with the provided recipes (e.g., pre-train/recipes/midtrain_v4_512.sh) to produce CoDA checkpoints (GCS or local storage).

Supervised Diffusion Fine-tuning (post-train/)

  1. Install prerequisites following post-train/LLaMA-Factory/README.md.
  2. Configure dataset metadata in post-train/LLaMA-Factory/data/dataset_info.json and diffusion arguments in post-train/LLaMA-Factory/examples/train_full/*.yaml.
  3. Execute ./run_sft.sh to fine-tune CoDA checkpoints with discrete denoising objectives.

Evaluation Pipelines (evaluation/)

  1. Choose a benchmark script such as evaluation/lm_eval/eval_mbpp_humaneval.sh.
  2. Update MODEL_DIR and diffusion parameters (diffusion_steps, temperature, top_p) to match the target checkpoint.
  3. Run the script to gather metrics; logs are stored locally for aggregation and reporting.

Benchmark πŸ“Š

Comparison of code-generation performance across standard and plus-enhanced benchmarks. Evalplus is computed as the mean pass@1 on enhanced variants. Bold marks results where CoDA produces the strongest diffusion-model performance.

ModelHumaneval InstructHumaneval PlusMBPP InstructMBPP PlusEvalplus
CoDA-Base29.323.835.246.034.9
CoDA-Instruct54.347.647.263.255.4
Dream-Base56.750.068.757.453.7
Dream-7B-Instruct57.953.768.356.154.9
LLaDA-8B-Instruct35.431.731.528.630.2
Qwen3-1.7B66.561.646.265.963.8
Qwen2.5-Coder-1.5B43.936.669.258.647.6
Qwen2.5-Coder-1.5B-Instruct70.766.569.259.462.3
Gemma-3-1B-it39.635.439.463.549.5
LLaMA-3.2-1B-Instruct35.431.124.453.742.4

Deployment Guide πŸ› οΈ

python3 -m venv .venv
source .venv/bin/activate

2 Install dependencies

pip install -r serving/requirements.txt

3 Export your Hugging Face token

export HF_TOKEN="hf_..."

4 Serve the model on cuda

bash serving/fast-api/start_server.sh

The server will listen on http://localhost:8000.

5 Interact with the served model

python serving/fast-api/chat_cli.py --base-url http://localhost:8000 --model Salesforce/CoDA-v0-Instruct

Optional flags:

  • --stream to stream tokens as they are generated
  • --show-meta to display latency and token usage

Generation hyperparameters (env vars)

You can customize generation with these environment variables (defaults in parentheses):

  • MAX_TOKENS (256)
  • TEMPERATURE (0.0)
  • TOP_P (unset)
  • TOP_K (unset)
  • STEPS (128)
  • ALG ("entropy")
  • ALG_TEMP (0.1)
  • BLOCK_LENGTH (32)

Example:

export MAX_TOKENS=512
export TEMPERATURE=0.7
export TOP_P=0.9
export STEPS=128
export ALG=entropy
export ALG_TEMP=0.1
export BLOCK_LENGTH=32
bash serving/fast-api/start_server.sh

Citation πŸ“š

@misc{coda2025,
  title={CoDA: Coding LM via Diffusion Adaptation},
  author={Chen, Haolin and Wang, Shiyu and Qin, Can and Pang, Bo and Liu, Zuxin and Qiu, Jielin and Zhang, Jianguo and Zhou, Yingbo and Chen, Zeyuan and Xu, Ran and Heinecke, Shelby and Savarese, Silvio and Xiong, Caiming and Wang, Huan and Yao, Weiran},
  year={2025},
  publisher={Salesforce AI Research}
}
coding
diffusion-models
llm
post-training
pre-training
sft

Languages

Python

98.8%

Shell

1.0%