datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Slakh2100-FLAC-Redux-Reducedgemini-flash-2.0-speech
🎙️ Gemini Flash 2.0 Speech Dataset
This is a high quality synthetic speech dataset generated by Gemini Flash 2.0 via the Multimodal Live API. It contains speech from 2 speakers - Puck (Male) and Kore (Female) in English.
🏅 #1 Trending Audio Dataset in Feb 2025
🏅 Used in training of Kokoro TTS and LLaSA 1B
〽️ Stats
Total number of audio files: 47,256*2 = 94512Total duration: 1023527.20seconds (284.31 hours)
Average duration: 10.83 seconds
Shortest file: 0.6… See the full description on the dataset page: https://huggingface.co/datasets/shb777/gemini-flash-2.0-speech.Audio-FLAN-Dataset
Audio-FLAN Dataset (Paper)
(the FULL audio files and jsonl files are still updating)
An Instruction-Tuning Dataset for Unified Audio Understanding and Generation Across Speech, Music, and Sound.
1. Dataset Structure
The Audio-FLAN-Dataset has the following directory structure:
Audio-FLAN-Dataset/
├── audio_files/
│ ├── audio/
│ │ └── 177_TAU_Urban_Acoustic_Scenes_2022/
│ │ └── 179_Audioset_for_Audio_Inpainting/
│ │ └── ...
│ ├── music/
│ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/Audio-FLAN-Dataset.musdb18-hq-flac
MUSDB18-HQ (FLAC Optimized)
Only the audio payload is converted to lossless PCM-16 FLAC. The original columns are preserved: audio, path, and instrument.
Source dataset
This dataset is derived from the original MUSDB18-HQ dataset.
The original dataset card and license are the authoritative references for the source audio and annotations.
Only the audio payload was transcoded to lossless PCM-16 FLAC; paths, instrument labels, and source track structure were… See the full description on the dataset page: https://huggingface.co/datasets/roro128/musdb18-hq-flac.wing-flap-noise-audio-exampleswavenet_flashback
Dataset Card for "wavenet_flashback"
https://cloud.google.com/text-to-speech/docs/reference/rest/v1/text/synthesize#AudioConfig
sv-SE-Wavenet-{voice}
https://spraakbanken.gu.se/resurser/flashback-dator
MAESTRO-v3.0-FLACdummy-flac-single-exampleethiopian-speech-flat
Ethio Speech Copus — Afrivoices Ethiopian
📌 Overview
The Ethio Speech Corpus dataset is a multilingual speech corpus containing audio–text pairs across five Ethiopian languages.
It is designed to support the development of speech-to-text technologies for low-resource languages.
This dataset is part of the Afrivoices initiative — a collaborative effort to create a large-scale ASR dataset for African languages.
The broader goal of the initiative is to collect 600 hours… See the full description on the dataset page: https://huggingface.co/datasets/badrex/ethiopian-speech-flat.kazakh-speech-dataset
Kazakh Speech Dataset
If you find this dataset helpful please press 'like' button
Summary
The Kazakh Speech Dataset is a large-scale open-source speech corpus for the Kazakh language. This dataset is designed to support the development of Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) systems for the Kazakh language.
Dataset Statistics
Total Audio Duration: ~726 hours
Language: Kazakh (kk)
Audio Format: FLAC
Sampling Rate: 16kHz
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Flamme-VRM/kazakh-speech-dataset.Wikipedia-FA-EN-DeepSeek-V4-Flash-0731
Wikipedia Persian to English — DeepSeek V4 Flash 0731
Rolling, machine-generated English translations of Persian Wikipedia articles
from Reza2kn/Wikipedia-EN-FA-Accessibility-Bridge, configuration
full_articles_fa_without_en. 129,816 translations are
currently published in 26 immutable Parquet shards.
The target release contains 129,816 translations;
five source rows have empty plain_text and are not translated. Shards are
published only after 5,000 complete, validated records… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/Wikipedia-FA-EN-DeepSeek-V4-Flash-0731.FLARE-1k-Unified-T2VAGemini-2.0-Flash-Aoede-VoiceFLARE-1k-Audio-T2VAGemini-2.0-Flash-Fenrir-VoiceGemini-2.0-Flash-Kore-Voiceghomala-spoken-bible
Ghomálá' Spoken New Testament — aligned audio + trilingual text
Part of the Lingo / NativeAI language-preservation project. This is
~20 hours of spoken Ghomálá' (Ghomala, ISO bbj; a Grassfields Bantu language of
West Cameroon) — recorded readings of the New Testament — aligned chapter-by-chapter
with parallel text in Ghomálá', French, and English.
Spoken-language data is exactly what oral-first Cameroonian languages lack, which makes
this a rare resource for building ASR, TTS… See the full description on the dataset page: https://huggingface.co/datasets/flagship-ai/ghomala-spoken-bible.FLARE-1k-Audio-T2VAgemini_flash_ttsfleurs-flac
FLEURS-FLAC
A losslessly FLAC-compressed version of Google's FLEURS dataset covering 102 languages.
Overview
This repository contains the Google FLEURS dataset repackaged into Parquet shards with PCM24 FLAC-compressed audio binaries.
Key points:
Audio streams are converted to FLAC (PCM24) with sample-level PCM verification against the source.
Sharded into ~500MB Parquet files per split for efficient I/O and streaming.
Covers all 102 languages from the original… See the full description on the dataset page: https://huggingface.co/datasets/roro128/fleurs-flac.flaviaGemini-2.0-Flash-Puck-VoiceFLARE-1k-Unified-T2VAGenderQA-gemini-1.5-flash
Dataset Card for "GenderQA-gemini-1.5-flash"
More Information needed
EmotionQA-gemini-1.5-flash-fix
Dataset Card for "EmotionQA-gemini-1.5-flash-fix"
More Information needed
LanguageQA-gemini-1.5-flash
Dataset Card for "LanguageQA-gemini-1.5-flash"
More Information needed
audio-arena-audio-flamingo-3-hfmusic-flamingo-dataset
Music Flamingo Distillation Dataset 2025
This dataset contains the Music Flamingo distillation data for training and evaluation.
Dataset Description
Music Flamingo is a model for music understanding and generation tasks. This dataset was originally hosted on Kaggle and has been migrated to HuggingFace for easier access and integration with the Hugging Face ecosystem.
Original Source
Originally available at:… See the full description on the dataset page: https://huggingface.co/datasets/beastLucifer/music-flamingo-dataset.datatalk-flac16kEmotionQA-gemini-1.5-flash
Dataset Card for "EmotionQA-gemini-1.5-flash"
More Information needed
