datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
biggest-ru-bookA bigger version of its5Q/bigger-ru-book, the smaller set being a subset of this one. Almost 1000 hours of high-quality audio.
ipapack_plus_train_3IWSLT.OfflineTaskantton-dataset
Antton Dataset (Synthetic)
This is a large-scale synthetic speech corpus designed for training and fine-tuning Basque Text-to-Speech (TTS) models. It consists of 99,996 audio files synthesized from the "Antton" voice model.
This dataset was generated by Itzune and serves as the primary source for training the itzune/antton-tts (Piper version) model.
Dataset Structure
Due to the large volume of data (approx. 100,000 files), the dataset is organized in the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/itzune/antton-dataset.INDICA
INDICA: An Audio Indic-Language Telecom Fraud Analysis Benchmark
Multilingual | Audio + Text | Benchmark for Fraud Detection
Overview
INDICA is a comprehensive benchmark for telecom fraud call analysis in Indic languages.It is built on the IndiF dataset, the first large-scale multilingual dataset for fraud detection in telecom conversations.
This benchmark enables research in:
Scenario Classification
Fraud Call Detection
Fraud-Type Classification… See the full description on the dataset page: https://huggingface.co/datasets/vikrant-vikram/INDICA.Multilingual_Speech_Dataset
Multilingual Speech Dataset
Paper: A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English
Repository: https://github.com/IS2AI/MultilingualASR
Description: This repository provides the dataset used in the paper "A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English". The paper focuses on training a single end-to-end (E2E) ASR model for Kazakh, Russian, and English, comparing monolingual and multilingual approaches… See the full description on the dataset page: https://huggingface.co/datasets/issai/Multilingual_Speech_Dataset.ipapack_plus_5ipapack_plus_3ipapack_plus_6bigger-ru-bookipapack_plus_1iter_1IMDAmaider-dataset
Maider Dataset (Synthetic)
This is a large-scale synthetic speech corpus designed for training and fine-tuning Basque Text-to-Speech (TTS) models. It consists of 99,996 audio files synthesized from the "Maider" voice model.
This dataset was generated by Itzune and serves as the primary source for training the itzune/maider-tts (Piper version) model.
Dataset Structure
Due to the large volume of data (approx. 100,000 files), the dataset is organized in the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/itzune/maider-dataset.ipapack_plus_4Speech-IFEvalaudio-qformerimproved-synthetic-vocal-burtsIALP-2026-data
IALP-2026: Whisper Open-Set Data-Selection — Query / Dev / Test Sets
Supporting data for the study "Whisper-Based Open-Set Data Selection for NSC
Adaptation." This repository holds the fixed target-query, validation, and
evaluation sets used across all experiments. Each part is a self-contained
.tar.gz.
All audio is 16 kHz mono. Each split ships with:
audio/ — audio files (FLAC, except GigaSpeech which is WAV PCM_16)
wav.scp — <utt_id> audio/<file> (Kaldi-style, relative paths)… See the full description on the dataset page: https://huggingface.co/datasets/pengyizhou/IALP-2026-data.improved_synthetic_vocal_burtsinternvideo2_dataChinese-Dialogue-180k-Instruct-Audiomultilingual-in-the-wildImagineV1-audioirescvttAudioX-IFcaps
[ICLR 2026] AudioX-IFcaps: Instruction-Following Audio Caption Dataset
AudioX-IFcaps (Instruction-Following) is a large-scale, high-quality multimodal dataset designed for training unified audio and music generation models. The dataset contains over 7 million samples with fine-grained, structured annotations that enable precise control over audio generation, including sound event categories, counts, temporal ordering, and timestamps.
📊 Dataset Statistics
General Audio:… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/AudioX-IFcaps.ICASSP2024-Acoustic_Scattering_AI-Noninvasive_Object_Classificationspt-br-tts-iasmin-qwen3
pt-br-tts-iasmin-qwen3
21957 clips PT-BR sintetizados com Qwen3-TTS-12Hz-1.7B. Voz Iasmin (voice-clone, is_iasmin=true, ~13957 clips) + vozes diversas CustomVoice (Ryan, Aiden, Vivian, Dylan, is_iasmin=false). WAV em tar shards (WebDataset); transcricao, voz, is_iasmin e sr em metadata.jsonl.
osu-beatmaps-duplicated
osu! Beatmaps Dataset (WebDataset)
A collection of ranked/loved osu! beatmaps with audio and chart data, in WebDataset format.
Dataset Variants
Variant
Audio Format
Description
original
MP3/OGG/WAV
Full quality original audio files
compressed
64kbps Mono Opus
Compressed audio for smaller download
from datasets import load_dataset
# Load original audio variant
ds = load_dataset("project-riz/osu-beatmaps", "original", streaming=True)
# Load compressed… See the full description on the dataset page: https://huggingface.co/datasets/IamXiangyu/osu-beatmaps-duplicated.Audio_speaker_needle_in_haystack
