datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
escucho-mucho-audio
Escucho Mucho — audio
Short Spanish speech clips (MP3, 24 kHz mono, ~64 kbps) used by the
Escucho Mucho
listening-practice app. Nothing here is original: the recordings are
re-encoded copies of public speech corpora, republished so the app can stream
them to a phone.
Each accent lives in its own folder; a clip's transcript, timings and
difficulty live in the app's own library index, not in this repo.
Folder
Source
Licence
co/
OpenSLR SLR72 — Colombian Spanish
CC BY-SA… See the full description on the dataset page: https://huggingface.co/datasets/likaili/escucho-mucho-audio.simple-escwa
🗣️ Simple-ESCWA: A simpler version of ESCWA-CS Corpus
The ESCWA-CS Corpus was collected over two days of meetings of the United Nations Economic and Social Commission for Western Asia (ESCWA) held in 2019.It contains intra-sentential code-switching between Arabic and English, with some speakers—particularly from Algeria, Tunisia, and Morocco—alternating between Arabic and French.
The dataset spans approximately 2.8 hours of speech, featuring dialectal Arabic and a Code Mixing Index… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/simple-escwa.esc-diagnostic-backup
ESC benchmark diagnostic dataset
Dataset Summary
As a part of ESC benchmark, we provide a small, 8h diagnostic dataset of in-domain validation data with newly annotated transcriptions. The audio data is sampled from each of the ESC validation sets, giving a range of different domains and speaking styles. The transcriptions are annotated according to a consistent style guide with two formats: normalised and un-normalised. The dataset is structured in the same way as the… See the full description on the dataset page: https://huggingface.co/datasets/esc-bench/esc-diagnostic-backup.
