datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aishell1-full-slr33
AISHELL-1 complete SLR33 archives
This is an independent, byte-preserving mirror of the complete AISHELL-1
archives published as OpenSLR SLR33. It is not affiliated with or endorsed by
AISHELL or OpenSLR.
Source: https://www.openslr.org/33/
data_aishell.tgz: complete speech data and transcripts
resource_aishell.tgz: lexicon, speaker information, and supplementary files
SHA256SUMS: checksums generated after download and gzip verification
The source is distributed under the… See the full description on the dataset page: https://huggingface.co/datasets/JazerJu/aishell1-full-slr33.navigation-corpus-speech-full-dagbani
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghana TTS Navigation Corpus — Dagbani
Synthetic speech dataset for navigation.
Structure
audio/ – all .wav audio files
text/ – matching .txt files with transcriptions
metadata.csv – full metadata table
ghana-female-twi-speech-asr-full-length
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Audio-text dataset with 76 pairs of Twi (Ghanaian language) speech data.
Structure
audio/ - WAV audio files ({len(pairs)} files)
text/ - Corresponding text transcripts ({len(pairs)} files)
dataset_manifest.json - Links audio to… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-female-twi-speech-asr-full-length.navigation-corpus-speech-full-twi
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghana TTS Navigation Corpus — Twi
Synthetic speech dataset for navigation.
Structure
audio/ – all .wav audio files
text/ – matching .txt files with transcriptions
metadata.csv – full metadata table
korean-full-duplex-synthetic-dataset-preview
Korean Full-Duplex Synthetic Dataset Preview
Overview
Public preview of a Korean full-duplex synthetic speech dataset. This
repository contains 100 conversations sampled from a corpus of 89,273
conversations (2,000.5 hours); it does not publish the full corpus audio.
Preview contents
100 conversation WAV files
data/representative.jsonl
24 kHz, mono, 16-bit PCM
Events: normal, barge_in, backchannel, cutoff_by_user
Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.tr-full-dataset
TR-Full_dataset
This is a merged speech dataset containing 41427 audio segments from 88 source datasets.
Dataset Information
Total Segments: 41427
Speakers: 222
Languages: tr
Emotions: neutral, angry, sad, happy
Original Datasets: 88
Dataset Structure
Each example contains:
audio: Audio file (WAV format, original sampling rate preserved)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/tr-full-dataset.navigation-corpus-speech-full-ewe
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghana TTS Navigation Corpus — Ewe
Synthetic speech dataset for navigation.
Structure
audio/ – all .wav audio files
text/ – matching .txt files with transcriptions
metadata.csv – full metadata table
VoxPopuli-Platinum-en-full
VoxPopuli-Platinum-en-full
Complete 142,066-row / ~404-hour English VoxPopuli Platinum dataset.
Reach out to data@trelis.com to purchase access or discuss a larger
custom-curation engagement.
Training Results
These results show why the Platinum labels matter. We compare the base model,
fine-tuning on raw VoxPopuli transcripts, and fine-tuning on this internally
filtered Platinum dataset. Evaluation uses the same english-spoken corpus WER
setup across four… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/VoxPopuli-Platinum-en-full.fleurs-farsi-fullopenslr61-es-ar-full
openslr61-es-ar-full
OpenSLR 61 (Crowdsourced high-quality Argentinian Spanish) consolidado COMPLETO con linaje.
Incluye male + female + weather messages argentinos.
Linaje (trazabilidad por sample)
source: siempre "openslr61"
subset: "main" (frases generales) o "weather" (mensajes de clima)
gender: "m" / "f"
speaker_id: ID anonimizado original del speaker
file_id: ID original del archivo OpenSLR
license: CC-BY-SA-4.0
Schema
campo
tipo… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/openslr61-es-ar-full.tr-full-dataset
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti Codyfederer tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: Codyfederer/tr-full-dataset
🔗 Derleyen Platform: VeriPazarı
TR-Full_dataset (Duygu Etiketli Türkçe Ses Veri Seti)
Bu veri seti, 88 kaynak veri setinden derlenmiş 41.427 ses segmentini (parçasını) içeren… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/tr-full-dataset.atco2-asr-atcosim
Dataset Card for "atco2-asr-atcosim"
This is a dataset constructed from two datasets: ATCO2-ASR and ATCOSIM.
It is divided into 80% train and 20% validation by selecting files randomly. Some of the files have additional information that is presented in the 'info' file.
adlam_fulfulde
Dataset Card for adlam fululde
This dataset contains 51 Pulaar speech recordings.
Credits and Acknowledgments
This work was produced by the Organisation pour la promotion de la langue Pulaar (Winden jangen ADLaM).
Contact Information
Website: www.windenjangen.org
Address: École Solokoure, Cimenterie, Conakry, Guinée
Emails: * windenjangen@windenjangen.org
aysha.sow12@gmail.com +224 622 15 40 75 +224 624463923
AMERICAN-ENGLISH-TRANSCRIBED-HIFI-FULL-DUPLEX-TWO-SPEAKER-CONVERSATIONAL-DATASET-SAMPLE
AMERICAN ENGLISH TRANSCRIBED HI-FI FULL-DUPLEX TWO-SPEAKER CONVERSATIONAL DATASET — SAMPLE
Overview
This open sample from Ocular AI contains four American English conversations between two people, with a separate audio track for each speaker and verbatim transcripts containing segment- and word-level timestamps.
The recordings capture conversational exchanges: repetitions, fillers, false starts, pauses, laughter, and audible breaths. Some conversations begin with… See the full description on the dataset page: https://huggingface.co/datasets/OcularAIInc/AMERICAN-ENGLISH-TRANSCRIBED-HIFI-FULL-DUPLEX-TWO-SPEAKER-CONVERSATIONAL-DATASET-SAMPLE.serendepify-gsl-asr-ak-waxal-gnlp-whisper-small-replay-fullft-v0.1
serendepify-gsl-asr-ak-waxal-gnlp-whisper-small-replay-fullft-v0.1
This is the first pushed Ghanaian Speech Lab ASR pipeline artifact. It is a
review artifact for the v0.1 Akan ASR pass, not a trained model checkpoint.
Expected future model repo:
teckedd/serendepify-gsl-asr-ak-waxal-gnlp-whisper-small-replay-fullft-v0.1
What This Artifact Contains
data/manifest.jsonl: harmonized Waxal + GhanaNLP manifest references.
reports/sanitize.json: sanitization report and… See the full description on the dataset page: https://huggingface.co/datasets/teckedd/serendepify-gsl-asr-ak-waxal-gnlp-whisper-small-replay-fullft-v0.1.lj_real_estate_deposition_full_case
LJ Real Estate Deposition – Full Case (Anonymized Audio + Transcript, EN)
Overview
This dataset contains a full real-world legal deposition in a real-estate / business context, provided as anonymized long-form audio + full court-style transcript.
It is designed for teams building and evaluating:
Automatic speech recognition (ASR) for long-form legal speech
Legal / real-estate conversation and dialogue models
Agent-style systems that need realistic, high-stakes… See the full description on the dataset page: https://huggingface.co/datasets/NorthAlabamaConsultants/lj_real_estate_deposition_full_case.tr-full-dataset
TR-Full_dataset
This is a merged speech dataset containing 41427 audio segments from 88 source datasets.
Dataset Information
Total Segments: 41427
Speakers: 222
Languages: tr
Emotions: neutral, angry, sad, happy
Original Datasets: 88
Dataset Structure
Each example contains:
audio: Audio file (WAV format, original sampling rate preserved)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion:… See the full description on the dataset page: https://huggingface.co/datasets/umutkkgz/tr-full-dataset.
