datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wavepulse-radio-raw-transcripts
WavePulse Radio Raw Transcripts
Dataset Summary
WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.ILRDF_Dict_Rukai
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ILRDF_Dict_Rukai
This is a noncanonical compatibility mirror. Use FormosanBank/ILRDF_Dicts for the complete canonical dataset and stable download contract.
This mirror… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ILRDF_Dict_Rukai.ILRDF_Dict_Bunun
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ILRDF_Dict_Bunun
This is a noncanonical compatibility mirror. Use FormosanBank/ILRDF_Dicts for the complete canonical dataset and stable download contract.
This mirror… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ILRDF_Dict_Bunun.ILRDF_Dict_Kavalan
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ILRDF_Dict_Kavalan
This is a noncanonical compatibility mirror. Use FormosanBank/ILRDF_Dicts for the complete canonical dataset and stable download contract.
This mirror… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ILRDF_Dict_Kavalan.ILRDF_Dict_Atayal
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ILRDF_Dict_Atayal
This is a noncanonical compatibility mirror. Use FormosanBank/ILRDF_Dicts for the complete canonical dataset and stable download contract.
This mirror… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ILRDF_Dict_Atayal.ILRDF_Dict_Puyuma
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ILRDF_Dict_Puyuma
This is a noncanonical compatibility mirror. Use FormosanBank/ILRDF_Dicts for the complete canonical dataset and stable download contract.
This mirror… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ILRDF_Dict_Puyuma.ILRDF_Dict_Yami
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ILRDF_Dict_Yami
This is a noncanonical compatibility mirror. Use FormosanBank/ILRDF_Dicts for the complete canonical dataset and stable download contract.
This mirror… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ILRDF_Dict_Yami.ILRDF_Dict_Paiwan
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ILRDF_Dict_Paiwan
This is a noncanonical compatibility mirror. Use FormosanBank/ILRDF_Dicts for the complete canonical dataset and stable download contract.
This mirror… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ILRDF_Dict_Paiwan.ILRDF_Dict_Saisiyat
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ILRDF_Dict_Saisiyat
This is a noncanonical compatibility mirror. Use FormosanBank/ILRDF_Dicts for the complete canonical dataset and stable download contract.
This mirror… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ILRDF_Dict_Saisiyat.ILRDF_Dict_Amis
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ILRDF_Dict_Amis
This is a noncanonical compatibility mirror. Use FormosanBank/ILRDF_Dicts for the complete canonical dataset and stable download contract.
This mirror… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ILRDF_Dict_Amis.ILRDF_Dict_Truku
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ILRDF_Dict_Truku
This is a noncanonical compatibility mirror. Use FormosanBank/ILRDF_Dicts for the complete canonical dataset and stable download contract.
This mirror… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ILRDF_Dict_Truku.ilrdf_dicts
Stage 3 資料集組裝與分層切分報告 (ILRDF Dicts)
語言總數: 16
資料集總筆數: 95,955 筆
資料集總時長: 126.59 小時 (455,736.6 秒)
切分策略: 確定性分層抽樣 (Seed=42)
排除重複音訊 (Deduplicated Audio): 2,008 筆 (3.03 小時, 文字不一致: 1,912 筆)
訓練集 (Train): 95,955 筆 (126.59 小時, 100.0%)
評估集 (Eval): 0 筆 (0.00 小時, 0.0%)
📋 各語言詳細組裝與切分統計
語言代碼
語言名稱
Train 筆數
Train 時長
Eval 筆數
Eval 時長
總筆數
總時長
Eval 佔比
ami-x-skl
秀姑巒阿美語
5,396
7.23 hrs
0
0.00 hrs
5,396
7.23 hrs
0.0%
bnn-x-isbk
郡群布農語
8,537
11.07 hrs
0
0.00 hrs
8,537
11.07 hrs… See the full description on the dataset page: https://huggingface.co/datasets/formospeech/ilrdf_dicts.ibo-dictigbo-dict is an Igbo text-audio dataset that includes the following:
25,500 single word audio recordings for each dialectal word variation
25,000 single Igbo sentence audio recording for each Igbo-English sentence pairing
Referenced in The IgboAPI Dataset: Empowering Igbo Language Technologies through Multi-dialectal Enrichment
safi-diction-sample
Safi Diction Sample
This dataset is a sample speech dataset containing short audio recordings paired with expected transcription text from respondents.
The dataset was created to show a sample of the type of data that can be collected with Safi's collection engine. Responses were collected remotely within the span of 12 hours.
This is just a sample for development or testing - contact hq@safidata.com for the full dataset or custom datasets, or visit https://www.safidata.com/.… See the full description on the dataset page: https://huggingface.co/datasets/martinturuta/safi-diction-sample.med-dictate
Med-Dictate — ASR Evaluation Dataset
An evaluation dataset released by Corti ApS alongside the Symphony for Speech Recognition white-paper. Medical notes dictated by Corti team members and one contractor, with their written consent, in English, French, and German. Built for benchmarking automatic speech recognition (ASR) and related NLP systems on medical-domain audio.
No real patient data. No PHI. No identifiable third-party content.
Languages: en, fr, de
How to… See the full description on the dataset page: https://huggingface.co/datasets/corti/med-dictate.igbo-dict-expansion-16khzBased on: https://huggingface.co/datasets/nkowaokwu/ibo-dict-expansion
The original audios were converted to 16kHz WAV.
Citation
If you want to cite this dataset you can use this:
@misc{igbo-dict-expansion-16khz,
title={Igbo dataset},
author={Jimenez, David},
howpublished={\url{https://huggingface.co/datasets/deepdml/igbo-dict-expansion-16khz}},
year={2025}
}
ILRDF_Dict_Saaroa
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ILRDF_Dict_Saaroa
This is a noncanonical compatibility mirror. Use FormosanBank/ILRDF_Dicts for the complete canonical dataset and stable download contract.
This mirror… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ILRDF_Dict_Saaroa.igbo-dict-expansionigbo-dict-16khzigbo-dict is an Igbo text-audio dataset that includes the following:
25,500 single word audio recordings for each dialectal word variation
25,000 single Igbo sentence audio recording for each Igbo-English sentence pairing
The original audios were converted to 16kHz WAV.
Referenced in The IgboAPI Dataset: Empowering Igbo Language Technologies through Multi-dialectal Enrichment
Citation
If you want to cite this dataset you can use this:
@misc{igbo-dict-16khz,
title={Igbo… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/igbo-dict-16khz.ILRDF_Dicts
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ILRDF_Dicts
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.
This is a canonical public audio dataset… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ILRDF_Dicts.ILRDF_Dict_Kanakanavu
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ILRDF_Dict_Kanakanavu
This is a noncanonical compatibility mirror. Use FormosanBank/ILRDF_Dicts for the complete canonical dataset and stable download contract.
This mirror… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ILRDF_Dict_Kanakanavu.igbo-dictAdditional_Yoruba_Datamed-dictate
Med-Dictate — ASR Evaluation Dataset
An evaluation dataset released by Corti ApS alongside the Symphony for Speech Recognition white-paper. Medical notes dictated by Corti team members and one contractor, with their written consent, in English, French, and German. Built for benchmarking automatic speech recognition (ASR) and related NLP systems on medical-domain audio.
No real patient data. No PHI. No identifiable third-party content.
Languages: en, fr, de
How to… See the full description on the dataset page: https://huggingface.co/datasets/MengjieChi/med-dictate.aura-phone-dictation-eval
Aura Phone Dictation Eval
Evaluation set of 365 progressive audio clips from 142 phone-number dictation sequences extracted from Aura Hindi/English call-center recordings.
This dataset is used to evaluate end-of-turn (EOT) detection models on structured phone-number dictation. Each sequence captures a caller dictating a 10-digit Indian mobile number across multiple speech segments. Progressive clips accumulate earlier segments plus trailing silence, ending with a final clip once… See the full description on the dataset page: https://huggingface.co/datasets/ananth-r-gnani/aura-phone-dictation-eval.ILRDF_Dict_Tsou
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ILRDF_Dict_Tsou
This is a noncanonical compatibility mirror. Use FormosanBank/ILRDF_Dicts for the complete canonical dataset and stable download contract.
This mirror… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ILRDF_Dict_Tsou.ILRDF_Dict_Thao
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ILRDF_Dict_Thao
This is a noncanonical compatibility mirror. Use FormosanBank/ILRDF_Dicts for the complete canonical dataset and stable download contract.
This mirror… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ILRDF_Dict_Thao.nigeria_ascentwarrungu-dictionaryThis dataset contains cleaned dictionary and grammar resources for the Warrungu language, compiled from structured CSV files for use in language revitalisation apps and AI tutors.
Files
Cleaned_Warrungu_Dictionary.csv
warrungu_flashcards.csv
warrungu_suffix_table.csv
/images/ (Warrungu flashcard images)
/audio/ (Warrungu flashcard audio)
License
Creative Commons Attribution 4.0 International (CC BY 4.0)
Contact
Maintained by the Warrungu project team.… See the full description on the dataset page: https://huggingface.co/datasets/warrungu/warrungu-dictionary.ikema_dict_asr
