datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
barranquenho-ipa-dict-synthetic
Barranquenho IPA Pronunciation Dictionary
The first and only IPA pronunciation dictionary of Barranquenho — the
Ibero-Romance contact variety spoken in Barrancos (Baixo Alentejo, Portugal),
a mixed system born of centuries of Portuguese–Spanish (Extremaduran /
Andalusian) contact on the raia. Every headword is written in the
Convenção Ortográfica do Barranquenho (2025) orthography and paired with a
broad-phonemic IPA transcription plus Portuguese and Spanish glosses.
This… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/barranquenho-ipa-dict-synthetic.med-dictate
Med-Dictate — ASR Evaluation Dataset
An evaluation dataset released by Corti ApS alongside the Symphony for Speech Recognition white-paper. Medical notes dictated by Corti team members and one contractor, with their written consent, in English, French, and German. Built for benchmarking automatic speech recognition (ASR) and related NLP systems on medical-domain audio.
No real patient data. No PHI. No identifiable third-party content.
Languages: en, fr, de
How to… See the full description on the dataset page: https://huggingface.co/datasets/corti/med-dictate.safi-diction-sample
Safi Diction Sample
This dataset is a sample speech dataset containing short audio recordings paired with expected transcription text from respondents.
The dataset was created to show a sample of the type of data that can be collected with Safi's collection engine. Responses were collected remotely within the span of 12 hours.
This is just a sample for development or testing - contact hq@safidata.com for the full dataset or custom datasets, or visit https://www.safidata.com/.… See the full description on the dataset page: https://huggingface.co/datasets/martinturuta/safi-diction-sample.igbo-dict-expansion-16khzBased on: https://huggingface.co/datasets/nkowaokwu/ibo-dict-expansion
The original audios were converted to 16kHz WAV.
Citation
If you want to cite this dataset you can use this:
@misc{igbo-dict-expansion-16khz,
title={Igbo dataset},
author={Jimenez, David},
howpublished={\url{https://huggingface.co/datasets/deepdml/igbo-dict-expansion-16khz}},
year={2025}
}
igbo-dict-expansionigbo-dict-16khzigbo-dict is an Igbo text-audio dataset that includes the following:
25,500 single word audio recordings for each dialectal word variation
25,000 single Igbo sentence audio recording for each Igbo-English sentence pairing
The original audios were converted to 16kHz WAV.
Referenced in The IgboAPI Dataset: Empowering Igbo Language Technologies through Multi-dialectal Enrichment
Citation
If you want to cite this dataset you can use this:
@misc{igbo-dict-16khz,
title={Igbo… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/igbo-dict-16khz.igbo-dictmed-dictate
Med-Dictate — ASR Evaluation Dataset
An evaluation dataset released by Corti ApS alongside the Symphony for Speech Recognition white-paper. Medical notes dictated by Corti team members and one contractor, with their written consent, in English, French, and German. Built for benchmarking automatic speech recognition (ASR) and related NLP systems on medical-domain audio.
No real patient data. No PHI. No identifiable third-party content.
Languages: en, fr, de
How to… See the full description on the dataset page: https://huggingface.co/datasets/MengjieChi/med-dictate.aura-phone-dictation-eval
Aura Phone Dictation Eval
Evaluation set of 365 progressive audio clips from 142 phone-number dictation sequences extracted from Aura Hindi/English call-center recordings.
This dataset is used to evaluate end-of-turn (EOT) detection models on structured phone-number dictation. Each sequence captures a caller dictating a 10-digit Indian mobile number across multiple speech segments. Progressive clips accumulate earlier segments plus trailing silence, ending with a final clip once… See the full description on the dataset page: https://huggingface.co/datasets/ananth-r-gnani/aura-phone-dictation-eval.Chinese-English-dictionary-resources
English and Chinese IPA Lexicons and Phoneme Sets
This repository provides English and Chinese IPA pronunciation lexicons and phoneme inventories collected from the vocabularies of multiple ASR corpora. The resources can be used for ASR, grapheme-to-phoneme conversion, pronunciation modeling, TTS, and related speech research.
Repository Structure
.
├── en/
│ ├── lexicon.txt
│ └── phone_list
└── zh/
├── lexicon.txt
└── phone_list
Source… See the full description on the dataset page: https://huggingface.co/datasets/maxwellziweiwei/Chinese-English-dictionary-resources.
