TuRTLe: A Unified Evaluation of LLMs for RTL Generation 🐢 (MLCAD 2025, ACM TODAES 2026)
See the code
TuRTLe is a framework to assess LLMs across key RTL generation tasks systematically. It integrates multiple existing benchmarks and automates the evaluation process, enabling a comprehensive assessment of LLM performance in syntax correctness, functional correctness, synthesis, PPA optimization, and exact line completion.
|
|
For more details about the framework, refer to the associated TuRTLe paper. For details on the NotSoTiny benchmark, see the corresponding paper.
Check the TuRTLe Leaderboard to know the best open-source models for each task.
Make sure you have installed TuRTLe and its dependencies. See Installation Guide for detailed setup instructions.
TuRTLe supports API-based inference which works out of the box with any OpenAI-compatible API (OpenRouter, OpenAI, Azure, etc.) with a Docker-based evaluation to run EDA tools locally.
$ export TURTLE_BASE_URL=https://openrouter.ai/api/v1
$ export TURTLE_API_KEY=sk-or-v1-...
$ uv run turtle/src/turtle.py --use-api \
--model mistralai/codestral-2508 \
--task notsotiny --shuttle tt06 \ # Can be tt06, tt07, tt08, tt09, tt10_ihp_25a, tt10_ihp_02, ttsky25a
--temperature 0.2 \
--max-tokens 131072 \
--top_p 0.95 \
--n_samples 1 \
--save_generations \
--save_generations_path './results/codestral-2508/nst-tt06.jsonl' \
--generation_only
Available tasks: rtllm, verilog_eval_rtl, verilog_eval_cc, verigen, rtlrepo
Evaluate the generated RTL designs using our bundled EDA tools (OpenLane, Verilator, Icarus Verilog):
$ docker run --rm -v $(pwd):/work -w /work ggcr0/turtle-eval:2.3.4 \
python3 turtle/src/turtle.py \
--task notsotiny \
--shuttle tt06 \
--model mistralai/codestral-2508 \
--n_samples 1 \
--load_generations_path ./results/codestral-2508/nst-tt06.jsonl
This will automatically pull the Docker image with all the EDA tooling and evaluate your designs for syntax, functionality, synthesis, and PPA metrics.
If you have access to a GPU cluster and want to run local inference with vLLM or perform multi-node inference, see LOCAL_INFERENCE.md for detailed instructions on using SLURM and Singularity.
The process to implement a benchmark is very similar to the one described by bigcode-evaluation-harness guide. Follow these steps:
turtle/tasks/template/new_task.py into turtle/tasks/ and rename it to the name of your benchmark <benchmark_name>.py._load_new_modules() and _create_extended_registry() methods within turtle/src/utils/task_updater.py.@inproceedings{garciagasulla2025turtleunifiedevaluationllms,
title={TuRTLe: A Unified Evaluation of LLMs for RTL Generation},
author={Dario Garcia-Gasulla and Gokcen Kestor and Emanuele Parisi and Miquel Albert\'i-Binimelis and Cristian Gutierrez and Razine Moundir Ghorab and Orlando Montenegro and Bernat Homs and Miquel Moreto},
booktitle = {Proceedings of the 2025 ACM/IEEE International Symposium on Machine Learning for CAD},
series = {MLCAD '25}
year={2025},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
location = {Santa Cruz, CA, USA},
url={https://arxiv.org/abs/2504.01986},
}
@article{10.1145/3831369,
author = {Alberti-Binimelis, Miquel and Gutierrez-Gomez, Cristian and Garcia-Gasulla, Dario and Parisi, Emanuele and Moundir Ghorab, Razine and Montenegro, Orlando and Homs, Bernat and Moreto, Miquel and Kestor, Gokcen},
title = {Revisiting TuRTLe: A Comprehensive Evaluation of LLMs for RTL Generation},
year = {2026},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
issn = {1084-4309},
url = {https://doi.org/10.1145/3831369},
doi = {10.1145/3831369},
note = {Just Accepted},
journal = {ACM Trans. Des. Autom. Electron. Syst.},
month = jul,
keywords = {Large Language Models (LLMs), Benchmarking, Computer-aided design (CAD)}
}
If you have any inquiries or wish to collaborate: hpai@bsc.es
This work was born as a fork of bigcode-evaluation-harness and vllm-code-harness, and has grown to its own framework for RTL code generation evaluation. We remain grateful to these projects.
We acknowledge the open-source EDA tools: Icarus Verilog, Verilator, Yosys, OpenROAD and LibreLane.
We also thank the authors of the benchmarks integrated in TuRTLe: VerilogEval, RTLLM, VGen, RTL-Repo and NotSoTiny.
Python
66.9%
Verilog
31.6%
Shell
1.5%
TuRTLe: A Unified Evaluation of LLMs for RTL Generation 🐢 (MLCAD 2025, ACM TODAES 2026)
See the code
TuRTLe is a framework to assess LLMs across key RTL generation tasks systematically. It integrates multiple existing benchmarks and automates the evaluation process, enabling a comprehensive assessment of LLM performance in syntax correctness, functional correctness, synthesis, PPA optimization, and exact line completion.
|
|
For more details about the framework, refer to the associated TuRTLe paper. For details on the NotSoTiny benchmark, see the corresponding paper.
Check the TuRTLe Leaderboard to know the best open-source models for each task.
Make sure you have installed TuRTLe and its dependencies. See Installation Guide for detailed setup instructions.
TuRTLe supports API-based inference which works out of the box with any OpenAI-compatible API (OpenRouter, OpenAI, Azure, etc.) with a Docker-based evaluation to run EDA tools locally.
$ export TURTLE_BASE_URL=https://openrouter.ai/api/v1
$ export TURTLE_API_KEY=sk-or-v1-...
$ uv run turtle/src/turtle.py --use-api \
--model mistralai/codestral-2508 \
--task notsotiny --shuttle tt06 \ # Can be tt06, tt07, tt08, tt09, tt10_ihp_25a, tt10_ihp_02, ttsky25a
--temperature 0.2 \
--max-tokens 131072 \
--top_p 0.95 \
--n_samples 1 \
--save_generations \
--save_generations_path './results/codestral-2508/nst-tt06.jsonl' \
--generation_only
Available tasks: rtllm, verilog_eval_rtl, verilog_eval_cc, verigen, rtlrepo
Evaluate the generated RTL designs using our bundled EDA tools (OpenLane, Verilator, Icarus Verilog):
$ docker run --rm -v $(pwd):/work -w /work ggcr0/turtle-eval:2.3.4 \
python3 turtle/src/turtle.py \
--task notsotiny \
--shuttle tt06 \
--model mistralai/codestral-2508 \
--n_samples 1 \
--load_generations_path ./results/codestral-2508/nst-tt06.jsonl
This will automatically pull the Docker image with all the EDA tooling and evaluate your designs for syntax, functionality, synthesis, and PPA metrics.
If you have access to a GPU cluster and want to run local inference with vLLM or perform multi-node inference, see LOCAL_INFERENCE.md for detailed instructions on using SLURM and Singularity.
The process to implement a benchmark is very similar to the one described by bigcode-evaluation-harness guide. Follow these steps:
turtle/tasks/template/new_task.py into turtle/tasks/ and rename it to the name of your benchmark <benchmark_name>.py._load_new_modules() and _create_extended_registry() methods within turtle/src/utils/task_updater.py.@inproceedings{garciagasulla2025turtleunifiedevaluationllms,
title={TuRTLe: A Unified Evaluation of LLMs for RTL Generation},
author={Dario Garcia-Gasulla and Gokcen Kestor and Emanuele Parisi and Miquel Albert\'i-Binimelis and Cristian Gutierrez and Razine Moundir Ghorab and Orlando Montenegro and Bernat Homs and Miquel Moreto},
booktitle = {Proceedings of the 2025 ACM/IEEE International Symposium on Machine Learning for CAD},
series = {MLCAD '25}
year={2025},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
location = {Santa Cruz, CA, USA},
url={https://arxiv.org/abs/2504.01986},
}
@article{10.1145/3831369,
author = {Alberti-Binimelis, Miquel and Gutierrez-Gomez, Cristian and Garcia-Gasulla, Dario and Parisi, Emanuele and Moundir Ghorab, Razine and Montenegro, Orlando and Homs, Bernat and Moreto, Miquel and Kestor, Gokcen},
title = {Revisiting TuRTLe: A Comprehensive Evaluation of LLMs for RTL Generation},
year = {2026},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
issn = {1084-4309},
url = {https://doi.org/10.1145/3831369},
doi = {10.1145/3831369},
note = {Just Accepted},
journal = {ACM Trans. Des. Autom. Electron. Syst.},
month = jul,
keywords = {Large Language Models (LLMs), Benchmarking, Computer-aided design (CAD)}
}
If you have any inquiries or wish to collaborate: hpai@bsc.es
This work was born as a fork of bigcode-evaluation-harness and vllm-code-harness, and has grown to its own framework for RTL code generation evaluation. We remain grateful to these projects.
We acknowledge the open-source EDA tools: Icarus Verilog, Verilator, Yosys, OpenROAD and LibreLane.
We also thank the authors of the benchmarks integrated in TuRTLe: VerilogEval, RTLLM, VGen, RTL-Repo and NotSoTiny.
Python
66.9%
Verilog
31.6%
Shell
1.5%