The Bangor Miami Corpus is a naturalistic Spanish-English code-switching speech dataset collected by Jon Russell Herring at Bangor University. It captures spontaneous bilingual conversations recorded in Miami, Florida, involving proficient Spanish-English bilinguals across multiple speaker groups.
| Total recordings | 56 |
| Total duration | ~32 h |
| Languages | English (en), Spanish (es) |
| Format | MP3 audio + plain-text transcript |
| Speaker groups | herring (16), maria (15), sastre (13), zeledon (12) |
Each row is one full recording session and contains:
| Column | Type | Description |
|---|---|---|
audio | Audio | Raw waveform decoded at original sampling rate |
file_name | string | Recording identifier (e.g. herring1, sastre3) |
transcript | string | Full conversation transcript, one utterance per line, cleaned of CHAT annotation markers |
spanish_share | float | Fraction of Spanish words in the recording (0 = all English, 1 = all Spanish), computed from word-level language-ID annotations in the corpus |
languages | list[string] | Always ["en", "es"] |
duration_s | float | Recording duration in seconds |
spanish_share is computed as:
spanish_share = count(langid == "spa") / count(langid != "999")
where langid comes from the word-level TSV annotations shipped with the corpus and 999 marks punctuation tokens.
The corpus spans from near-monolingual English (< 2 %) to near-monolingual Spanish (> 95 %), with a mean of ~34 % Spanish words.
defaultAll 56 recordings (~32 h total).
mixedA ~2.5 h subset of 5 recordings restricted to genuinely mixed conversations where 20%–80% of content words are Spanish. Recordings are drawn from all four speaker groups and cover the full 0.2–0.8 Spanish-share range.
| Speaker | File | Duration | Spanish share |
|---|---|---|---|
| herring | herring17 | ~30 min | 0.37 |
| maria | maria20 | ~32 min | 0.37 |
| sastre | sastre8 | ~33 min | 0.39 |
| zeledon | zeledon4 | ~22 min | 0.43 |
| zeledon | zeledon14 | ~33 min | 0.79 |
The original corpus was collected and transcribed at Bangor University. If you use this dataset, please cite the original work:
@misc{bangor_miami,
author = {Deuchar, Margaret and Davies, Peredur and Herring, Jon Russell
and Parafita Couto, Maria Carmen and Carter, Diana},
title = {Building bilingual corpora},
booktitle = {Bilingualism: Basic principles and beyond},
editor = {Thomas, Enlli and Mennen, Ineke},
year = {2014},
publisher = {Multilingual Matters},
address = {Bristol}
}
The corpus is available under CC BY-SA 3.0.
The Bangor Miami Corpus is a naturalistic Spanish-English code-switching speech dataset collected by Jon Russell Herring at Bangor University. It captures spontaneous bilingual conversations recorded in Miami, Florida, involving proficient Spanish-English bilinguals across multiple speaker groups.
| Total recordings | 56 |
| Total duration | ~32 h |
| Languages | English (en), Spanish (es) |
| Format | MP3 audio + plain-text transcript |
| Speaker groups | herring (16), maria (15), sastre (13), zeledon (12) |
Each row is one full recording session and contains:
| Column | Type | Description |
|---|---|---|
audio | Audio | Raw waveform decoded at original sampling rate |
file_name | string | Recording identifier (e.g. herring1, sastre3) |
transcript | string | Full conversation transcript, one utterance per line, cleaned of CHAT annotation markers |
spanish_share | float | Fraction of Spanish words in the recording (0 = all English, 1 = all Spanish), computed from word-level language-ID annotations in the corpus |
languages | list[string] | Always ["en", "es"] |
duration_s | float | Recording duration in seconds |
spanish_share is computed as:
spanish_share = count(langid == "spa") / count(langid != "999")
where langid comes from the word-level TSV annotations shipped with the corpus and 999 marks punctuation tokens.
The corpus spans from near-monolingual English (< 2 %) to near-monolingual Spanish (> 95 %), with a mean of ~34 % Spanish words.
defaultAll 56 recordings (~32 h total).
mixedA ~2.5 h subset of 5 recordings restricted to genuinely mixed conversations where 20%–80% of content words are Spanish. Recordings are drawn from all four speaker groups and cover the full 0.2–0.8 Spanish-share range.
| Speaker | File | Duration | Spanish share |
|---|---|---|---|
| herring | herring17 | ~30 min | 0.37 |
| maria | maria20 | ~32 min | 0.37 |
| sastre | sastre8 | ~33 min | 0.39 |
| zeledon | zeledon4 | ~22 min | 0.43 |
| zeledon | zeledon14 | ~33 min | 0.79 |
The original corpus was collected and transcribed at Bangor University. If you use this dataset, please cite the original work:
@misc{bangor_miami,
author = {Deuchar, Margaret and Davies, Peredur and Herring, Jon Russell
and Parafita Couto, Maria Carmen and Carter, Diana},
title = {Building bilingual corpora},
booktitle = {Bilingualism: Basic principles and beyond},
editor = {Thomas, Enlli and Mennen, Ineke},
year = {2014},
publisher = {Multilingual Matters},
address = {Bristol}
}
The corpus is available under CC BY-SA 3.0.