datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gemini-flash-2.0-speech
🎙️ Gemini Flash 2.0 Speech Dataset
This is a high quality synthetic speech dataset generated by Gemini Flash 2.0 via the Multimodal Live API. It contains speech from 2 speakers - Puck (Male) and Kore (Female) in English.
🏅 #1 Trending Audio Dataset in Feb 2025
🏅 Used in training of Kokoro TTS and LLaSA 1B
〽️ Stats
Total number of audio files: 47,256*2 = 94512Total duration: 1023527.20seconds (284.31 hours)
Average duration: 10.83 seconds
Shortest file: 0.6… See the full description on the dataset page: https://huggingface.co/datasets/shb777/gemini-flash-2.0-speech.wavenet_flashback
Dataset Card for "wavenet_flashback"
https://cloud.google.com/text-to-speech/docs/reference/rest/v1/text/synthesize#AudioConfig
sv-SE-Wavenet-{voice}
https://spraakbanken.gu.se/resurser/flashback-dator
ethiopian-speech-flat
Ethio Speech Copus — Afrivoices Ethiopian
📌 Overview
The Ethio Speech Corpus dataset is a multilingual speech corpus containing audio–text pairs across five Ethiopian languages.
It is designed to support the development of speech-to-text technologies for low-resource languages.
This dataset is part of the Afrivoices initiative — a collaborative effort to create a large-scale ASR dataset for African languages.
The broader goal of the initiative is to collect 600 hours… See the full description on the dataset page: https://huggingface.co/datasets/badrex/ethiopian-speech-flat.ghomala-spoken-bible
Ghomálá' Spoken New Testament — aligned audio + trilingual text
Part of the Lingo / NativeAI language-preservation project. This is
~20 hours of spoken Ghomálá' (Ghomala, ISO bbj; a Grassfields Bantu language of
West Cameroon) — recorded readings of the New Testament — aligned chapter-by-chapter
with parallel text in Ghomálá', French, and English.
Spoken-language data is exactly what oral-first Cameroonian languages lack, which makes
this a rare resource for building ASR, TTS… See the full description on the dataset page: https://huggingface.co/datasets/flagship-ai/ghomala-spoken-bible.fleurs-flac
FLEURS-FLAC
A losslessly FLAC-compressed version of Google's FLEURS dataset covering 102 languages.
Overview
This repository contains the Google FLEURS dataset repackaged into Parquet shards with PCM24 FLAC-compressed audio binaries.
Key points:
Audio streams are converted to FLAC (PCM24) with sample-level PCM verification against the source.
Sharded into ~500MB Parquet files per split for efficient I/O and streaming.
Covers all 102 languages from the original… See the full description on the dataset page: https://huggingface.co/datasets/roro128/fleurs-flac.
