datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wiktionary-ipa-audio-en
English Wiktionary IPA + audio
English pronunciation rows extracted from the structured Kaikki/Wiktextract
English dump, restricted to English entries with both IPA and a playable
Wikimedia Commons recording. The dataset contains one row per pronunciation
and recording pairing; an audio recording can therefore occur in more than
one row when Wiktionary associates it with multiple IPA or entry records.
Fields
The audio column is created by the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/wiktionary-ipa-audio-en.wiktionary-ipa-audio
Wiktionary IPA Audio Mirror
This is a materialized audio mirror of mostol/wiktionary-ipa, published so
users do not need to individually download Wikimedia Commons URLs.
This initial snapshot is partial and contains the successfully downloaded
deduplicated rows available at publication time. metadata.csv maps each
WAV file to its IPA transcription and original source URL. Audio is 16 kHz,
mono, PCM WAV.
The original source URLs are retained for provenance and attribution. Please… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/wiktionary-ipa-audio.
