datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-asr-leaderboard-multilingual-datasets
ASR Leaderboard Datasets
This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS).
How to Load
To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>.
from datasets import load_dataset
# Load the FLEURS dataset for Bulgarian
fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg")
print(fleurs_bg)
# Load the… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-multilingual-datasets.opendata-iisys-hui
HUI-Audio-Corpus-German Dataset
Overview
The HUI-Audio-Corpus-German is a high-quality Text-To-Speech (TTS) dataset developed by researchers at the Institute of Information Systems (IISYS). This dataset is designed to facilitate the development and training of TTS applications, particularly in the German language. The associated research paper can be found here.
Dataset Contents
The dataset comprises recordings from multiple speakers, with the five most… See the full description on the dataset page: https://huggingface.co/datasets/Paradoxia/opendata-iisys-hui.open-large-bengali-asr-data
Open Large Bengali ASR Data
This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps).
Datasets:
commonvoice
openslr
madasr
shrutilipi
flerus
kathbath
indictts
ucla
gali
open-large-bengali-asr-data
Open Large Bengali ASR Data
This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps).
Datasets:
commonvoice
bn-asr-mega-open-dataKirundi_Open_Speech_Dataset
🇧🇮 Kirundi Open Speech & Text Dataset
Building the first large-scale, open-source speech and text dataset for Kirundi
🚀 Get Started • 📊 Dataset • 🎯 Roadmap • 🫱🏿🫲🏾 Community
🌍 About This Project
Kirundi is spoken by over 12 million people, yet it remains a low-resource language largely ignored by modern AI systems. We're changing that.
This community-driven initiative aims to create the first comprehensive, open-source speech and text dataset for… See the full description on the dataset page: https://huggingface.co/datasets/Ijwi-ry-Ikirundi-AI/Kirundi_Open_Speech_Dataset.WanJuanSiLu-Multimodal-5Languages
WanJuan·SiLu Multimodal Multilingual Corpus
🌏Dataset Introduction
The newly upgraded "Wanjuan·Silk Road Multimodal Corpus" brings the following three core improvements:
The number of languages has been significantly expanded: Based on the five open-source languages of "Wanjuan·Silk Road", namely Arabic, Russian, Korean, Vietnamese, and Thai, "Wanjuan·Silk Road Multimodal" has added three scarce corpus data of Serbian, Hungarian, and Czech, and uses the above… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/WanJuanSiLu-Multimodal-5Languages.cs_open_data_asropen_data_asropen-music-dataset-demo
Dataset Card for "open-music-dataset-demo"
More Information needed
open-asr-leaderboard-multilingual-datasets
Open ASR Leaderboard Armenian Test Datasets
This private repository holds leaderboard-compatible Armenian test
configurations while their integration is being validated.
Configurations
fleurs_hy
Source: google/fleurs,
configuration hy_am, test split
Reviewed reference changes:
Metric-AI/fleurs-corrections,
test split
932 recordings; all 314 reviewed corrections were matched to the original
source transcript and applied
mcv_hy… See the full description on the dataset page: https://huggingface.co/datasets/Metric-AI/open-asr-leaderboard-multilingual-datasets.tendances-audio-video-barometre
Tendances audio-vidéo - Baromètre
[!NOTE]
Ce jeu de données Hugging Face est vide. Cette carte sert seulement à référencer le jeu de données Tendances audio-vidéo - Baromètre qui est disponible à l'adresse https://www.data.gouv.fr/datasets/6836e0dc1baaf48fbb8b1851
Description
L’Arcom publie les données du volet quantitatif de son Baromètre Tendances audio-vidéo. L’étude est conduite auprès d’un échantillon représentatif de Français âgés de 15 ans et plus.
Elle vise à… See the full description on the dataset page: https://huggingface.co/datasets/french-open-data/tendances-audio-video-barometre.open-bambara-asr-datasethallym_AI_OpenDataset
Hallym Adult and Child Speech Dataset
This dataset contains speech recordings and transcriptions collected from adult and child speakers for AI-based speech and language research.
Dataset Overview
Total Records: 2,714
Speakers: 49 (adult: 25, child: 24)
Groups: adult, child
File Format: WAV (audio) + TXT (transcription)
Speaker Statistics
Group
Count
Gender
Age Range
Adult
25명
남/여
50~78세
Child
24명
남/여
3~8세
Dataset Fields… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/hallym_AI_OpenDataset.WanJuanSiLu-Multimodal-3Languages
WanJuan·SiLu Multimodal Multilingual Corpus
🌏Dataset Introduction
The newly upgraded "Wanjuan·Silk Road Multimodal Corpus" brings the following three core improvements:
The number of languages has been significantly expanded: Based on the five open-source languages of "Wanjuan·Silk Road", namely Arabic, Russian, Korean, Vietnamese, and Thai, "Wanjuan·Silk Road Multimodal" has added three scarce corpus data of Serbian, Hungarian, and Czech, and uses the above eight key… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/WanJuanSiLu-Multimodal-3Languages.
