datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic_speech_corpus
Dataset Card for Arabic Speech Corpus
Dataset Summary
This Speech corpus has been developed as part of PhD work carried out by Nawar Halabi at the University of Southampton. The corpus was recorded in south Levantine Arabic (Damascian accent) using a professional studio. Synthesized speech as an output using this corpus has produced a high quality, natural voice.
Supported Tasks and Leaderboards
[Needs More Information]
Languages
The audio is in… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/arabic_speech_corpus.Korean-Japanese-Code-Switching-Speech
Korean-Japanese-Code-Switching-Speech
This dataset contains Korean-Japanese code-switching speech recordings with sentence-level transcriptions. It was introduced in the paper Towards Truly Multilingual ASR: Generalizing Code-Switching ASR to Unseen Language Pairs.
Since there is an extremely small amount of Korean-Japanese code-switching data available, this dataset was designed to be used as a small-scale evaluation dataset.
The dataset consists of code-switching recordings… See the full description on the dataset page: https://huggingface.co/datasets/thetaone-ai/Korean-Japanese-Code-Switching-Speech.speech-mendeley-pa
Credit - https://data.mendeley.com/datasets/sdbc8f5b77/2
Kirundi_Open_Speech_Dataset
🇧🇮 Kirundi Open Speech & Text Dataset
Building the first large-scale, open-source speech and text dataset for Kirundi
🚀 Get Started • 📊 Dataset • 🎯 Roadmap • 🫱🏿🫲🏾 Community
🌍 About This Project
Kirundi is spoken by over 12 million people, yet it remains a low-resource language largely ignored by modern AI systems. We're changing that.
This community-driven initiative aims to create the first comprehensive, open-source speech and text dataset for… See the full description on the dataset page: https://huggingface.co/datasets/Ijwi-ry-Ikirundi-AI/Kirundi_Open_Speech_Dataset.speech-pa
Dataset Card for "speech-pa"
More Information needed
m-ailabs_speech_dataset_fr\
The M-AILABS Speech Dataset is the first large dataset that we are providing free-of-charge, freely usable as training data for speech recognition and speech synthesis.
Most of the data is based on LibriVox and Project Gutenberg. The training data consist of nearly thousand hours of audio and the text-files in prepared format.
A transcription is provided for each clip. Clips vary in length from 1 to 20 seconds and have a total length of approximately shown in the list (and in the respective info.txt-files) below.
The texts were published between 1884 and 1964, and are in the public domain. The audio was recorded by the LibriVox project and is also in the public domain – except for Ukrainian.
Ukrainian audio was kindly provided either by Nash Format or Gwara Media for machine learning purposes only (please check the data info.txt files for details).bloom-speechBloom-speech is a dataset of text aligned speech from bloomlibrary.org. This dataset contains over 50 languages including many low-resource languages. This dataset should be useful for training and/or testing speech-to-text or text-to-speech/ASR models.mandarin-speech-samples
Mandarin Speech Samples
This sample shows Mandarin Chinese speech with clip-level metadata and preview transcripts. It is meant to help buyers review language fit, recording quality, and sample structure before scoping a larger delivery.
What This Shows
Mandarin speech audio with consistent metadata
Clip-level transcript fields for content review
Language and format signals for procurement review
Dataset Specifications
Field
Value… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/mandarin-speech-samples.bengali-multi-speaker-speech-samples
Bengali Speech: Multi-Speaker Samples
This sample shows Bengali multi-speaker speech with aligned ground-truth transcripts. It is meant to help buyers review conversational structure, speaker overlap, transcript quality, and audio consistency before scoping a larger delivery.
What This Shows
Multi-speaker Bengali speech with transcript alignment
Conversation-style audio rather than isolated prompt reading
Metadata that distinguishes language, format, and speaker… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bengali-multi-speaker-speech-samples.tamil-speech-samples
Tamil Speech Samples
This sample shows Tamil speech with ground-truth transcripts and consistent audio metadata. It is meant to help buyers review language fit, transcript quality, and capture format before scoping a larger delivery.
What This Shows
Tamil speech recordings with paired transcripts
Ground-truth labels at the clip level
Format metadata for review and delivery planning
Dataset Specifications
Field
Value
Modality
Audio… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/tamil-speech-samples.korean-speech-samples
Korean Speech Samples
This sample shows Korean contributor speech in a consistent audio format. It is meant to help buyers review recording quality, language coverage, and metadata structure before scoping a larger delivery.
What This Shows
Korean speech recordings from contributor collection workflows
Clip-level metadata for format and review context
Ground-truth transcripts for understanding sample content
Dataset Specifications
Field… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/korean-speech-samples.cantonese-speech-samples
Cantonese Speech Samples
This sample shows Cantonese speech with native transcript metadata. It is meant to help buyers review dialect fit, recording quality, and sample structure before requesting broader coverage.
What This Shows
Cantonese speech audio with paired transcript metadata
Language-specific metadata for review and delivery planning
A compact preview of the available sample structure
Dataset Specifications
Field
Value… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/cantonese-speech-samples.french-speech-samples
French Speech Samples
This sample shows French contributor speech paired with source transcripts. It is meant to help buyers review recording quality, transcript alignment, and metadata structure before scoping a larger delivery.
What This Shows
French single-speaker recordings
Transcript alignment from the source dataset
Clip-level metadata for format and review context
Dataset Specifications
Field
Value
Modality
Audio
Language… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/french-speech-samples.urdu-speech-samples
Urdu Speech Samples
This sample shows Urdu speech with transcript alignment and simple audio metadata. It is meant to help buyers review language fit and capture quality before requesting a larger sample or production delivery.
What This Shows
Urdu speech audio with paired text
A compact view of transcript and metadata structure
Audio format signals for procurement review
Dataset Specifications
Field
Value
Modality
Audio
Language
Urdu… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/urdu-speech-samples.telugu-speech-samples
Telugu Speech Samples
This sample shows Telugu speech with ground-truth transcripts and consistent audio metadata. It is meant to help buyers review language fit, transcript quality, and capture format before scoping a larger delivery.
What This Shows
Telugu speech recordings with paired transcripts
Ground-truth labels at the clip level
Format metadata for review and delivery planning
Dataset Specifications
Field
Value
Modality
Audio… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/telugu-speech-samples.henan-speech-samples
Henan Speech Samples
This sample shows Henan Chinese speech with native transcript metadata. It is meant to help buyers review regional speech fit, recording quality, and sample structure before requesting broader coverage.
What This Shows
Henan Chinese speech audio with paired transcript metadata
Regional language metadata for review and delivery planning
A compact preview of the available sample structure
Dataset Specifications
Field
Value… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/henan-speech-samples.hindi-speech-samples
Hindi Speech Samples
This sample shows Hindi contributor speech paired with validated text. It is meant to help buyers review spoken content, transcript alignment, and audio format consistency before scoping a larger delivery.
What This Shows
Single-speaker Hindi recordings from contributor collection workflows
Ground-truth transcript alignment at the clip level
Audio metadata suitable for evaluating format and capture consistency
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/hindi-speech-samples.spanish-speech-samples
Spanish Speech Samples
This sample shows Spanish contributor speech in a consistent audio format. It is meant to help buyers review recording quality, language coverage, and metadata structure before scoping a larger delivery.
What This Shows
Spanish speech samples with clip-level review metadata
Clip-level metadata for format and review context
Ground-truth transcripts for understanding sample content
Dataset Specifications
Field
Value… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/spanish-speech-samples.bengali-speech-samples
Bengali Speech Samples
This sample shows Bengali read and conversational speech with paired transcripts. It is meant to help buyers review spoken content, transcript alignment, and audio consistency before scoping a larger delivery.
What This Shows
Bengali speech across read and conversational styles
Clip-level transcript alignment
Audio metadata that supports format and quality review
Dataset Specifications
Field
Value
Modality
Audio… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bengali-speech-samples.ultravox-indic-speech
Equal Indic Speech Dataset 1
A large-scale multilingual speech dataset for English and Hindi, containing 4.2 million audio-transcript pairs
Dataset Overview
Total Samples: 4,226,934 audio-transcript pairs
Languages: English and Hindi
Total Duration: ~X,XXX hours of audio
Sample Rate: Various (preserved from source)
Dataset (Train + Validation)
English Train: 1,290,610 samples
English Validation: 368,745 samples
Hindi Train: 1,997,008 samples
Hindi Validation: 570,571… See the full description on the dataset page: https://huggingface.co/datasets/equal-ai/ultravox-indic-speech.Reazon-Speech
Reazon Speech v2 dataset mirror
Original Dataset Source
Hugging Face Dataset Page: reazon-research/reazonspeech
Project Page: Reazon Research
License
This dataset is a mirror of the original Reazon Speech v2 dataset, but on 🤗 server (so may be faster). This dataset is licensed under the CDLA-Sharing-1.0. The original dataset comes with the following restriction:
TO USE THIS DATASET, YOU MUST AGREE THAT YOU WILL USE THE DATASET SOLELY FOR THE PURPOSE OF… See the full description on the dataset page: https://huggingface.co/datasets/samson-ailabs/Reazon-Speech.
