This dataset is a collection of natural language (English) instructions and corresponding Bash commands for the task of natural language to Bash translation (NL2SH).
This dataset contains a test set of 600 manually verified instruction-command pairs, and a training set of 40,639 unverified pairs for the development and benchmarking of machine translation models. The creation of NL2SH-ALFA was motivated by the need for larger, more accurate NL2SH datasets. The associated InterCode-ALFA benchmark uses the NL2SH-ALFA test set. For more information, please refer to the paper.
Note, the config parameter, NOT the split parameter, selects the train/test data.
from datasets import load_dataset
train_dataset = load_dataset("westenfelder/NL2SH-ALFA", "train", split="train")
test_dataset = load_dataset("westenfelder/NL2SH-ALFA", "test", split="train")
This dataset is intended for training and evaluating NL2SH models.
This dataset is not intended for natural languages other than English, scripting languages other than Bash, nor multi-line Bash scripts.
The training set contains two columns:
nl: string - natural language instructionbash: string - Bash commandThe test set contains four columns:
nl: string - natural language instructionbash: string - Bash commandbash2: string - Bash command (alternative)difficulty: int - difficulty level (0, 1, 2) corresponding to (easy, medium, hard)Both sets are unordered.
The NL2SH-ALFA dataset was created to increase the amount of NL2SH training data and to address errors in the test sets of previous datasets.
The dataset was produced by combining, deduplicating and filtering multiple datasets from previous work. Additionally, it includes instruction-command pairs scraped from the tldr-pages. Please refer to Section 4.1 of the paper for more information about data collection, processing and filtering.
Source datasets:

find.BibTeX:
@inproceedings{westenfelder-etal-2025-llm,
title = "{LLM}-Supported Natural Language to Bash Translation",
author = "Westenfelder, Finnian and
Hemberg, Erik and
Moskal, Stephen and
O{'}Reilly, Una-May and
Chiricescu, Silviu",
editor = "Chiruzzo, Luis and
Ritter, Alan and
Wang, Lu",
booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
month = apr,
year = "2025",
address = "Albuquerque, New Mexico",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.naacl-long.555/",
pages = "11135--11147",
ISBN = "979-8-89176-189-6"
Finn Westenfelder
Please email finnw@mit.edu or make a pull request.
This dataset is a collection of natural language (English) instructions and corresponding Bash commands for the task of natural language to Bash translation (NL2SH).
This dataset contains a test set of 600 manually verified instruction-command pairs, and a training set of 40,639 unverified pairs for the development and benchmarking of machine translation models. The creation of NL2SH-ALFA was motivated by the need for larger, more accurate NL2SH datasets. The associated InterCode-ALFA benchmark uses the NL2SH-ALFA test set. For more information, please refer to the paper.
Note, the config parameter, NOT the split parameter, selects the train/test data.
from datasets import load_dataset
train_dataset = load_dataset("westenfelder/NL2SH-ALFA", "train", split="train")
test_dataset = load_dataset("westenfelder/NL2SH-ALFA", "test", split="train")
This dataset is intended for training and evaluating NL2SH models.
This dataset is not intended for natural languages other than English, scripting languages other than Bash, nor multi-line Bash scripts.
The training set contains two columns:
nl: string - natural language instructionbash: string - Bash commandThe test set contains four columns:
nl: string - natural language instructionbash: string - Bash commandbash2: string - Bash command (alternative)difficulty: int - difficulty level (0, 1, 2) corresponding to (easy, medium, hard)Both sets are unordered.
The NL2SH-ALFA dataset was created to increase the amount of NL2SH training data and to address errors in the test sets of previous datasets.
The dataset was produced by combining, deduplicating and filtering multiple datasets from previous work. Additionally, it includes instruction-command pairs scraped from the tldr-pages. Please refer to Section 4.1 of the paper for more information about data collection, processing and filtering.
Source datasets:

find.BibTeX:
@inproceedings{westenfelder-etal-2025-llm,
title = "{LLM}-Supported Natural Language to Bash Translation",
author = "Westenfelder, Finnian and
Hemberg, Erik and
Moskal, Stephen and
O{'}Reilly, Una-May and
Chiricescu, Silviu",
editor = "Chiruzzo, Luis and
Ritter, Alan and
Wang, Lu",
booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
month = apr,
year = "2025",
address = "Albuquerque, New Mexico",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.naacl-long.555/",
pages = "11135--11147",
ISBN = "979-8-89176-189-6"
Finn Westenfelder
Please email finnw@mit.edu or make a pull request.