datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
portuguese-unified-pronunciation-lexicon
Portuguese Unified Pronunciation Lexicon
A flat, single-row-per-pronunciation dataset merging Portuguese IPA transcriptions from three authoritative sources. Each row is a word × region × POS tuple with both broad phonemic (ipa_broad) and narrow phonetic (ipa_narrow) transcriptions normalized across sources.
Source
Words
Convention
Description
Infopédia (Porto Editora)
102,685
Broad phonemic
European Portuguese dictionary IPA
Wiktionary (pt.wiktionary.org)
15,720… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-unified-pronunciation-lexicon.english-pronunciation-audiowiktionary_pronunciations-finalPronunciation-dictionary-malayalam
Malayalam Pronunciation Dictionary
This Dataset has an alternate name of Malayalam Phonetic Lexicon. It is curated from the original source here
It gives Phonemic transcription of Malayalam words in IPA format.
Dataset Details
Dataset Description
This is a collection of Malayalam words and their pronunciation described in IPA format. The pronunciations has been automatically generated using [Mlphon]
(https://pypi.org/project/mlphon/) Python library.
Curated… See the full description on the dataset page: https://huggingface.co/datasets/kavyamanohar/Pronunciation-dictionary-malayalam.wiktionary_pronunciations-backupPronunciation-boldvoice
Pronunciation Assessment Dataset (BoldVoice + speechocean762)
Dataset for fine-tuning multimodal models on English pronunciation assessment.
Overview
Source
Samples
Audio Duration
Description
BoldVoice
38,182
10-20s
Non-native English learners, BoldVoice API annotations
speechocean762
5,000
1.6-20s
Public dataset, 5-expert scored, Mandarin speakers
Total
43,182
Schema
Column
Type
Description
audio
Audio (16kHz mono)
Speech… See the full description on the dataset page: https://huggingface.co/datasets/aigc-x/Pronunciation-boldvoice.wiktionary_pronunciations
Dataset Card for "wiktionary_pronunciations"
More Information needed
saskia_may1_pronunciationmandarin_pronunciation_qaUyghur-Consonant-Vowel-Combo-Pronunciationsusername_pronunciationdrugs-and-substances-wiktionary-pronunciations
TigreGotico/drugs-and-substances-wiktionary-pronunciations
Rich entity dataset scraped by metadatarr
scraper wiktionary_pronunciations.
Part of the drugs-and-substances collection.
Rows: 1,200
Fields
term
language
ipa
wikitext_excerpt
wiktionary_url
source_wiktionary
Source
Generated by scrapers/wiktionary_pronunciations.py. See the metadatarr repo for the full
pipeline and scraper source code.
Saskia_may6_pronunciation_allpersian_word_vowels_pronunciations
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/developer-ninja/persian_word_vowels_pronunciations.parle-french-pronunciation-datasets
Parle French Pronunciation Datasets
Version v2026.08.11 · DOI 10.5281/zenodo.21932586
This first-party open-data release documents two bounded learning structures used by Parle: French Pronunciation, an iPhone and iPad App from Tingnova Inc. for English-speaking complete beginners and A0–A2 learners.
Ting Dong, the creator of Parle, maintains this dataset for Tingnova Inc. This repository is product documentation and open educational data. It is not an independent review… See the full description on the dataset page: https://huggingface.co/datasets/Levindong/parle-french-pronunciation-datasets.Saskia_may5_pronunciation_allPronunciation_Dictionaries_Alsatian_Dialects
[!NOTE]
Dataset origin: https://zenodo.org/records/1174214
Description
This dataset contains a collection of pronunciation dictionaries which were manually transcribed using the X-SAMPA transcription system. The transcriptions were performed based on audio recordings available on the following websites :
OLCA: http://www.olcalsace.org/fr/lexiques
Elsässich Web diktionnair: http://www.ami-hebdo.com/elsadico/index.php
The dataset was produced in the context of the RESTAURE project… See the full description on the dataset page: https://huggingface.co/datasets/IAlsace/Pronunciation_Dictionaries_Alsatian_Dialects.wikipedia-for-pronunciations-tempwiktionary_pronunciations-testpronunciationenglish-pronunciationpronunciation_control
Dataset Card for "pronunciation_control"
More Information needed
Technical_Terms_With_Pronunciations_and_AudioProSign-SignLanguage-Pronunciation
