datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TTS-Italian
TTS-Italian
A high-quality Italian speech dataset for text-to-speech and automatic speech recognition.
Data Sources
Derived from LibriVox Italian — volunteer-read Italian public domain audiobooks hosted on archive.org.
Books: 23 Italian-language audiobooks (Dante, Pirandello, Verga, De Amicis, Collodi, Pascoli, etc.)
License: Public Domain
Processing: Standardized to 24kHz mono, WhisperX transcription (large-v3) with word-level alignment, segmented at word boundaries… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Italian.multilingual_librispeech_italian_phoneme
Multilingual LibriSpeech Italian Phoneme
Dataset Summary
This dataset is a curated version of the Italian subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme).
The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into Italian acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_italian_phoneme.Italian-Speech-Dataset
Italian Dataset Metadata
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Italian (it)
🏷️ Tags
Audio, Speech, Speech Recognition, Machine, Machine Learning, ML
📦 Size Category
n < 1K
accenti_italiani
🤌 Accenti Italiani
Accenti Italiani is a dataset designed to evaluate Automatic Speech Recognition (ASR) models under challenging real-world conditions.It was created to test model robustness on strong regional Italian accents together with noisy audio, making the task deliberately difficult.
Overview
Format: WAV
Sampling rate: 16000 Hz
How to load
from datasets import load_dataset
from IPython.display import Audio, display
ds =… See the full description on the dataset page: https://huggingface.co/datasets/lopozz/accenti_italiani.italian-speech-recognition-dataset
Italian Telephone Dialogues Dataset - 499 Hours
The dataset provides 499 hours of annotated telephone dialogues from 676 native speakers in Italy. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of telephone… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/italian-speech-recognition-dataset.YodaLingua-Italian
YodaLingua-Italian
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Italian portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
58,500 audio–transcription pairs
Total duration
161 hours
Speakers
2,319 distinct speakers
Audio format
MP3 • mono • 24 kHz • 16-bit… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Italian.ItalianConversationalDataset
Italian Multimodal Conversations Dataset
The largest Italian-language conversational dataset combining real chat dialogues, audio transcripts, images, and verified purchase transactions — built for training LLMs, chatbots, and multimodal AI models.
📊 Dataset Overview
Asset
Count
Details
💬 Conversations
1,000,000+
Tagged by topic, scenario, participants
🎙️ Audio Messages
9,889
With full transcriptions
🖼️ Images
5,321
Linked to conversation context
💳… See the full description on the dataset page: https://huggingface.co/datasets/ciaocicciquantum/ItalianConversationalDataset.
