datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
twi-health-asr-gemini-500hrs
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Health Speech Dataset Gemini (500 hours)
A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken
languages, sourced from publicly available video content on health and wellness.
Created by… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-health-asr-gemini-500hrs.gemini-flash-2.0-speech
🎙️ Gemini Flash 2.0 Speech Dataset
This is a high quality synthetic speech dataset generated by Gemini Flash 2.0 via the Multimodal Live API. It contains speech from 2 speakers - Puck (Male) and Kore (Female) in English.
🏅 #1 Trending Audio Dataset in Feb 2025
🏅 Used in training of Kokoro TTS and LLaSA 1B
〽️ Stats
Total number of audio files: 47,256*2 = 94512Total duration: 1023527.20seconds (284.31 hours)
Average duration: 10.83 seconds
Shortest file: 0.6… See the full description on the dataset page: https://huggingface.co/datasets/shb777/gemini-flash-2.0-speech.twi-health-asr-gemini-500hrs
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Health Speech Dataset Gemini (500 hours)
A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken
languages, sourced from publicly available video content on health and wellness.
Created by… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/twi-health-asr-gemini-500hrs.twi-health-asr-gemini-500hrs-ipa
Twi Health Speech — Audio, Transcript and IPA
Twi health-domain speech with both a written transcript and an IPA phoneme sequence read off the audio by ASR. Built from ghananlpcommunity/twi-health-asr-gemini-500hrs by adding the IPA column.
from datasets import load_dataset
ds = load_dataset("ghananlpcommunity/twi-health-asr-gemini-500hrs-ipa", split="train")
ds[0]["audio"] # decoded waveform, 16 kHz
ds[0]["transcription"] # transcript
ds[0]["ipa"]… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/twi-health-asr-gemini-500hrs-ipa.wpp_pav_transcrito_gemini
🎤 Transcrições WhatsApp - Google Gemini 2.0 Flash
Este dataset contém transcrições de mensagens de áudio do WhatsApp geradas usando Google Gemini 2.0 Flash.
📋 Descrição
Origem: Mensagens de áudio do WhatsApp em português brasileiro
Modelo: Google Gemini 2.0 Flash
Preço: $0.075/1M tokens
Total de amostras: 198
Formato de áudio: WAV (16kHz)
Idioma: Português brasileiro
Modelo de IA multimodal avançado da Google com capacidades de análise contextual de áudio.… See the full description on the dataset page: https://huggingface.co/datasets/BernardoAI/wpp_pav_transcrito_gemini.target-words-geminitts
Target-Word (TW) Evaluation Set
Synthetic speech clips for 101 rare drug terms, intended for evaluation only — measuring how
well an ASR system recognises rare / out-of-vocabulary medical vocabulary (target-word WER / CER /
recall). Each clip reads a real DailyMed sentence containing one target drug name, synthesised with
Google Gemini TTS across multiple voices. This is the frozen evaluation set from the master's thesis
"Audio-free lexical adaptation of Whisper's decoder"… See the full description on the dataset page: https://huggingface.co/datasets/aharalambieva/target-words-geminitts.
