datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
evals-speech-recognition-cy-en
Welsh ASR Model Evaluation Transcription Dataset
This resource compiles the output transcriptions from multiple Welsh Automatic Speech Recognition (ASR) models across several test sets.
The data is structured hierarchically:
Splits delineate the individual test sets.
Configs within each split detail the performance (transcriptions) of a specific ASR model and its version on that set.
Metrics Results
model
test
task
wer
cer… See the full description on the dataset page: https://huggingface.co/datasets/techiaith/evals-speech-recognition-cy-en.CASIA_speech_emotion_recognitionevals-speech-recognition-cy-en-2606Quran_speech_recognition_kaggleThis dataset can be found in Kaggle
speech-emotion-recognition
Speech Emotion Recognition
Dataset comprises 30,000+ audio recordings featuring 4 distinct emotions: euphoria, joy, sadness, and surprise. This extensive collection is designed for research in emotion recognition, focusing on the nuances of emotional speech and the subtleties of speech signals as individuals vocally express their feelings.
By utilizing this dataset, researchers and developers can enhance their understanding of sentiment analysis and improve automatic speech… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/speech-emotion-recognition.portuguese-speech-recognition-dataset
Portuguese Speech Dataset for recognition task
Dataset comprises 10+ hours of telephone dialogues in Portuguese, collected from 10+ native speakers across various topics and domains. It is a valuable resource for advancing speech recognition technology.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech recognition (ASR) systems, transcribing audio, and natural language processing (NLP). - Get the data
The… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/portuguese-speech-recognition-dataset.spanish-speech-recognition-dataset
Spanish Speech Dataset for recognition task
Dataset comprises 10 hours of telephone dialogues in Spanish, collected from 10 native speakers across various topics and domains. It is a valuable resource for advancing speech recognition technology.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech recognition (ASR) systems, transcribing audio, and natural language processing (NLP). - Get the data
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/spanish-speech-recognition-dataset.american-speech-recognition-dataset
American Speech Dataset for recognition task
Dataset comprises 1,136 hours of telephone dialogues in American, collected from 1,416 native speakers across various topics and domains, achieving an impressive 95% Sentence Accuracy Rate. It is designed for research in automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in natural language processing (NLP), speech recognition, and machine… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/american-speech-recognition-dataset.ssi-speech-emotion-recognition
Dataset Card for SSI: Speech Emotion Recognition - Stapes AI
Dataset Details
Dataset Format for Audio Files
This is the format for the audio files in the dataset. We'll open-source the dataset soon.
Gender
M - Male
F - Female
Age Group
CH - Child (0-12)
TE - Teenager (13-19)
AD - Adult (20-60)
SE - Senior (60+)
UNK - Unknown
Utterance Type
SEN: Sentence
WOR: Word
PHR: Phrase
Sentence
DFA: "Don't Forget A… See the full description on the dataset page: https://huggingface.co/datasets/stapesai/ssi-speech-emotion-recognition.russian-speech-recognition-dataset
Russian Speech Dataset for recognition task
Dataset comprises 338 hours of telephone dialogues in Russian, collected from 460 native speakers across various topics and domains, with an impressive 98% Word Accuracy Rate. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/russian-speech-recognition-dataset.speech-recognition-data
Audio Speech Recognition Dataset
Dataset Card
Dataset Summary
This dataset comprises high-quality audio recordings of a standardized 50-word paragraph, designed for training and evaluating automatic speech recognition (ASR) models. Each recording captures the text: "AI is transforming our world, from voice assistants to self-driving cars. Good AI needs quality data. Your voice helps train speech recognition. Speak clearly, naturally, without noise. Be yourself.… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/speech-recognition-data.slovenian-speech-recognition
Slovenian Speech Dataset
Dataset comprises 10+ hours of audio recordings featuring 20+ speakers engaged in telephone dialogues in the Slovenian language. It contains speech data designed for training robust language models and automatic speech recognition systems in real-world conversational scenarios.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in natural language processing (NLP), speech recognition, and machine… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/slovenian-speech-recognition.arabic-speech-recognition
Arabic Speech Dataset
Dataset comprises over 10 hours of audio featuring 20+ native speakers engaged in telephone-quality dialogues in the Arabic language. It contains high-quality speech data designed for training robust language models and automatic speech recognition systems in real-world conversational scenarios.
By utilizing this dataset, developers and researchers can advance their work in automatic speech recognition and improve recognition systems. - Get the data
The… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/arabic-speech-recognition.japanese-speech-recognition-dataset
Japanese Speech Dataset for recognition task
Dataset comprises 10+ hours of telephone dialogues in Japanese, collected from 10 native speakers across various topics and domains. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/japanese-speech-recognition-dataset.speech-emotion-recognition-datasetThe audio dataset consists of a collection of texts spoken with four distinct
emotions. These texts are spoken in English and represent four different
emotional states: **euphoria, joy, sadness and surprise**.
Each audio clip captures the tone, intonation, and nuances of speech as
individuals convey their emotions through their voice.
The dataset includes a diverse range of speakers, ensuring variability in age,
gender, and cultural backgrounds*, allowing for a more comprehensive
representation of the emotional spectrum.
The dataset is labeled and organized based on the emotion expressed in each
audio sample, making it a valuable resource for emotion recognition and
analysis. Researchers and developers can utilize this dataset to train and
evaluate machine learning models and algorithms, aiming to accurately
recognize and classify emotions in speech.vietnamese-speech-recognition
Vietnamese Speech Dataset
Dataset comprises 10+ hours of telephone dialogues in Vietnamese, collected from 20 native speakers across various topics and domains. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in transcribing audio, and natural… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/vietnamese-speech-recognition.evals-speech-recognition-cy-en-2601
model
test
task
wer
cer
DewiBrynJones/whisper-large-v2-ft-cy-2601
cymen-arfor/lleisiau-arfor
transcribe
39.0266
18.9881
techiaith/kaldi-cy-2601
cymen-arfor/lleisiau-arfor
transcribe
54.421
27.4451
DewiBrynJones/whisper-large-v2-ft-cy-2601
techiaith/banc-trawsgrifiadau-bangor
transcribe
31.1593
12.5359
techiaith/kaldi-cy-2601
techiaith/banc-trawsgrifiadau-bangor
transcribe
45.2049
20.1329
DewiBrynJones/whisper-large-v2-ft-cy-2601
techiaith/commonvoice-23-0-cy
transcribe… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/evals-speech-recognition-cy-en-2601.CASIA_speech_emotion_recognition_preloadrussian-speech-recognition-dataset
Russian Speech Recognition Dataset - 1,000+ Hours Call Center
1,000+ hours of real-world Russian call center audio with transcripts. Train speech recognition, sentiment analysis, and customer support AI models on authentic telephone conversations
Dataset Summary
Key Features
✅ 1,000+ hours of inbound & outbound calls✅ 100% Russian telephone conversations✅ Real-world audio - no synthetic data✅ Full transcripts in Russian and in English
Full… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/russian-speech-recognition-dataset.Speech_recognition_dataevals-speech-recognition-cy-en-2511
model
test
wer
cer
DewiBrynJones/whisper-large-v3-ft-btb-cv-cvad-ca-cy-2511
cymen-arfor/lleisiau-arfor
31.4418
12.6665
DewiBrynJones/whisper-large-v3-ft-btb-cv-cvad-ca-wlga-cy-2511
cymen-arfor/lleisiau-arfor
29.3326
11.3554
DewiBrynJones/whisper-large-v2-ft-btb-cv-cvad-ca-wlga-cy-2511
cymen-arfor/lleisiau-arfor
28.1715
10.9305
DewiBrynJones/whisper-large-v3-ft-btb-cv-cvad-ca-cy-2511
techiaith/banc-trawsgrifiadau-bangor
27.686
9.6299… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/evals-speech-recognition-cy-en-2511.speech-recognition-congolese-languages
Speech Recognition Datasets for Congolese Languages
Dataset Details
Dataset Description
This dataset contains two new benchmark corpora designed for low-resource languages spoken in the Democratic Republic of the Congo: The Lingala Read Speech Corpus LRSC, with 4.3 hours of labelled audio, and the Congolese Speech Radio Corpus CSRC, which offers 741 hours of unlabeled audio spanning four significant low-resource languages of the region (Lingala, Tshiluba… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/speech-recognition-congolese-languages.korean-speech-recognition
Korean Speech Dataset
Dataset comprises 10+ hours of audio recordings from 20+ speakers, featuring telephone-quality speech data from native korean speakers. It provides a diverse collection of spoken language for automatic speech recognition tasks and serves as essential training data for model training in NLP and speech detection research.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech recognition… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/korean-speech-recognition.au.2.speech_recognition
Dataset Card for "au.2.speech_recognition"
More Information needed
british-english-speech-recognition-dataset
British English Speech Dataset for recognition task
Dataset comprises 200 hours of high-quality audio recordings featuring 310 speakers, achieving an impressive 95% Sentence Accuracy Rate. This extensive collection of speech data is designed for NLP tasks such as speech recognition, dialogue systems, and language understanding.
By utilizing this dataset, developers and researchers can advance their work in automatic speech recognition and improve recognition systems. - Get the… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/british-english-speech-recognition-dataset.arabic-speech-recognition
Arabic Speech Dataset - 10+ hours
Dataset comprises over 10 hours of audio featuring 20+ native speakers engaged in telephone-quality dialogues in the Arabic language. It contains high-quality speech data designed for training robust language models and automatic speech recognition systems in real-world conversational scenarios. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of telephone dialogues in Slovenian for training NLP… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/arabic-speech-recognition.french-speech-recognition-dataset
French Speech Dataset for recognition task
Dataset comprises 547 hours of telephone dialogues in French, collected from 964 native speakers across various topics and domains, with an impressive 98% Word Accuracy Rate. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/french-speech-recognition-dataset.hindi-speech-recognition-dataset
Hindi Speech Dataset for recognition task
Dataset comprises 760 hours of telephone dialogues in Hindi, collected from 1,000+ native speakers across various topics and domains. This dataset boasts an impressive 95% sentence accuracy rate, making it a valuable resource for advancing speech recognition technology.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech recognition (ASR) systems, transcribing audio… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/hindi-speech-recognition-dataset.russian-speech-recognition-dataset
Russian Telephone Dialogues Dataset - 338 Hours
The Russian speech dataset includes 338 hours of telephone dialogues in Russian from 460 native speakers, offering high-quality audio recordings with detailed annotations (text, speaker ID, gender, age) to support speech recognition systems, natural language processing, and deep learning models for building accurate Russian dialogue and audio datasets. - Get the data
Dataset characteristics:
Characteristic
Data… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/russian-speech-recognition-dataset.hindi-speech-recognition-dataset
Hindi Telephone Dialogues Dataset - 760 Hours
Dataset comprises 760 hours of high-quality audio recordings from 1,000+ native Hindi speakers, featuring telephone dialogues across diverse topics and domains. With a 95% sentence accuracy rate, this essential dataset is ideal for training and evaluating Hindi speech recognition systems. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of telephone dialogues in Hindi for training… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/hindi-speech-recognition-dataset.
