datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
biggest-ru-bookA bigger version of its5Q/bigger-ru-book, the smaller set being a subset of this one. Almost 1000 hours of high-quality audio.
yodas-ja000
YODAS Japanese (ja000)
Japanese manual caption subset of the YODAS dataset, repackaged for easier use.
Source
Original dataset: espnet/yodas (ja000 config)
Paper: YODAS: YouTube-Oriented Dataset for Audio and Speech
License: CC BY 3.0
Citation
If you use this dataset, please cite the original YODAS paper:
fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/ItzmeNishh/fleurs.it_youtube_uzbek_speech_dataset
IT Uzbek Speech Dataset
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language and with some english to better generalization. The data was collected from publicly available videos on YouTube related to the Information Technology (IT) field. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Mohir Dev YouTube channel (respect to the team for… See the full description on the dataset page: https://huggingface.co/datasets/islomov/it_youtube_uzbek_speech_dataset.omnivoice-it
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
10,000
Total
20,000
TTS-Italian
TTS-Italian
A high-quality Italian speech dataset for text-to-speech and automatic speech recognition.
Data Sources
Derived from LibriVox Italian — volunteer-read Italian public domain audiobooks hosted on archive.org.
Books: 23 Italian-language audiobooks (Dante, Pirandello, Verga, De Amicis, Collodi, Pascoli, etc.)
License: Public Domain
Processing: Standardized to 24kHz mono, WhisperX transcription (large-v3) with word-level alignment, segmented at word boundaries… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Italian.antton-dataset
Antton Dataset (Synthetic)
This is a large-scale synthetic speech corpus designed for training and fine-tuning Basque Text-to-Speech (TTS) models. It consists of 99,996 audio files synthesized from the "Antton" voice model.
This dataset was generated by Itzune and serves as the primary source for training the itzune/antton-tts (Piper version) model.
Dataset Structure
Due to the large volume of data (approx. 100,000 files), the dataset is organized in the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/itzune/antton-dataset.it_youtube_uzbek_speech_dataset
IT Uzbek Speech Dataset
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language and with some english to better generalization. The data was collected from publicly available videos on YouTube related to the Information Technology (IT) field. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Mohir Dev YouTube channel (respect to the team for… See the full description on the dataset page: https://huggingface.co/datasets/BoburAmirov/it_youtube_uzbek_speech_dataset.ScreenTalk-XS
🎬 ScreenTalk-XS: Sample Speech Dataset from Screen Content 🖥️
📢 What is ScreenTalk-XS?
ScreenTalk-XS is a high-quality transcribed speech dataset containing 10k speech samples from diverse screen content.It is designed for automatic speech recognition (ASR), natural language processing (NLP), and conversational AI research.
✅ This dataset is freely available for research and educational use.🔹 If you need a larger dataset with more diverse speech samples… See the full description on the dataset page: https://huggingface.co/datasets/Itbanque/ScreenTalk-XS.multilingual_librispeech_italian_phoneme
Multilingual LibriSpeech Italian Phoneme
Dataset Summary
This dataset is a curated version of the Italian subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme).
The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into Italian acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_italian_phoneme.maider-dataset
Maider Dataset (Synthetic)
This is a large-scale synthetic speech corpus designed for training and fine-tuning Basque Text-to-Speech (TTS) models. It consists of 99,996 audio files synthesized from the "Maider" voice model.
This dataset was generated by Itzune and serves as the primary source for training the itzune/maider-tts (Piper version) model.
Dataset Structure
Due to the large volume of data (approx. 100,000 files), the dataset is organized in the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/itzune/maider-dataset.Italian-Speech-Dataset
Italian Dataset Metadata
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Italian (it)
🏷️ Tags
Audio, Speech, Speech Recognition, Machine, Machine Learning, ML
📦 Size Category
n < 1K
it_youtube_uzbek_speech_dataset
IT Uzbek Speech Dataset
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language and with some english to better generalization. The data was collected from publicly available videos on YouTube related to the Information Technology (IT) field. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Mohir Dev YouTube channel (respect to the team for… See the full description on the dataset page: https://huggingface.co/datasets/azimislom/it_youtube_uzbek_speech_dataset.accenti_italiani
🤌 Accenti Italiani
Accenti Italiani is a dataset designed to evaluate Automatic Speech Recognition (ASR) models under challenging real-world conditions.It was created to test model robustness on strong regional Italian accents together with noisy audio, making the task deliberately difficult.
Overview
Format: WAV
Sampling rate: 16000 Hz
How to load
from datasets import load_dataset
from IPython.display import Audio, display
ds =… See the full description on the dataset page: https://huggingface.co/datasets/lopozz/accenti_italiani.zwesui-grupa-5-it-ai
Wykorzystanie ASR do transkrypcji polskich nagrań o tematyce AI
Korpus do ewaluacji systemów ASR języka polskiego stworzony w ramach warsztatów
Ewaluacja Systemów Rozpoznawania Mowy (UAM WMI, edycja 2026, zespół 5).
Zbiór powstał jako część kursu - publikujemy go publicznie, żeby inni badacze
polskiego ASR mogli z niego korzystać i porównywać wyniki na wspólnym benchmarku.
Cel i pytania badawcze
Cel główny:
Porównanie jakości 3 systemów ASR dla spontanicznej… See the full description on the dataset page: https://huggingface.co/datasets/slapekm/zwesui-grupa-5-it-ai.ScreenTalk
🏢 ScreenTalk: Full Dataset
📌 Overview
ScreenTalk is a structured transcription dataset designed to improve automatic speech recognition (ASR) and natural language processing (NLP) models. It contains transcriptions from diverse screen content, capturing natural dialogues, different speech styles, and realistic conversational patterns.
The full version of ScreenTalk provides access to the complete dataset, covering multiple languages, genres, and a vast range of… See the full description on the dataset page: https://huggingface.co/datasets/Itbanque/ScreenTalk.YodaLingua-Italian
YodaLingua-Italian
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Italian portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
58,500 audio–transcription pairs
Total duration
161 hours
Speakers
2,319 distinct speakers
Audio format
MP3 • mono • 24 kHz • 16-bit… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Italian.
