The NST dataset was created to support the development of automatic speech recognition (ASR) and speech-to-text dictation systems in Norwegian. The dataset provides high-quality, annotated audio recordings for linguistic research and ASR system training. Originally developed by Nordisk Språkteknologi Holding AS (NST), the dataset was later acquired and preserved by a consortium of Norwegian institutions, including University of Oslo, University of Bergen, and the National Library of Norway.
Each example in the dataset includes:
The full dataset contains:
What are the data types?
No direct personal identifiers are included, but demographic metadata such as age, sex, and region is present.
The recordings were originally made by NST using multiple speakers under controlled environments with both close and distant microphone setups.
Speakers belong to varying demographics:
What preprocessing was done?
What labeling is provided?
What are the intended uses?
What are the out-of-scope uses?
The dataset is distributed under the Apache 2.0 License.
The dataset is hosted by NbAiLab on Hugging Face and curated by the National Library of Norway. Users are encouraged to provide feedback or contribute improved versions by contacting: sprakbanken@nb.no
28 commits
1 commits
The NST dataset was created to support the development of automatic speech recognition (ASR) and speech-to-text dictation systems in Norwegian. The dataset provides high-quality, annotated audio recordings for linguistic research and ASR system training. Originally developed by Nordisk Språkteknologi Holding AS (NST), the dataset was later acquired and preserved by a consortium of Norwegian institutions, including University of Oslo, University of Bergen, and the National Library of Norway.
Each example in the dataset includes:
The full dataset contains:
What are the data types?
No direct personal identifiers are included, but demographic metadata such as age, sex, and region is present.
The recordings were originally made by NST using multiple speakers under controlled environments with both close and distant microphone setups.
Speakers belong to varying demographics:
What preprocessing was done?
What labeling is provided?
What are the intended uses?
What are the out-of-scope uses?
The dataset is distributed under the Apache 2.0 License.
The dataset is hosted by NbAiLab on Hugging Face and curated by the National Library of Norway. Users are encouraged to provide feedback or contribute improved versions by contacting: sprakbanken@nb.no
28 commits
1 commits