CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ghanaopenai /twi_multispeaker_audio_transcribed Twi Multispeaker Audio Transcribed Dataset Overview The Twi Multispeaker Audio Transcribed dataset is a collection of speech recordings and their transcriptions in Asante Twi, a widely spoken dialect of the Akan language in Ghana. The dataset is designed for training and evaluating automatic speech recognition (ASR) models and other natural language processing (NLP) applications. Dataset Details Source: The dataset is derived from the Financial… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi_multispeaker_audio_transcribed.audioautomatic-speech-recognition10K<n<100K0 likes324 downloads2y agoHugging Face02ghanaopenai /akuapem_multispeaker_audio_transcribed Akuapem Multispeaker Audio Transcribed Dataset Overview The Akuapem Multispeaker Audio Transcribed dataset is a collection of speech recordings and their transcriptions in Akuapem Twi, a widely spoken dialect of the Akan language in Ghana. The dataset is designed for training and evaluating automatic speech recognition (ASR) models and other natural language processing (NLP) applications. Dataset Details Source: The dataset is derived from the… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/akuapem_multispeaker_audio_transcribed.audioautomatic-speech-recognition10K<n<100K1 likes269 downloads2y agoHugging Face03trentmkelly /wwii_audio_transcribed WWII Audio with Transcripts 993 World War II-era recordings from the Internet Archive WWII audio collection, with machine-generated transcripts: 1944: 558 recordings 1945: 435 recordings Total: approximately 220 hours of audio; 6.78 GB including alternate audio formats, transcripts, and archive images. The dataset viewer pairs each recording with playable audio and its full transcript. Transcripts were generated with Microsoft MAI Transcribe 2 and may contain errors or be… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/wwii_audio_transcribed.audioautomatic-speech-recognitionn<1K0 likes69 downloads3d agoHugging Face04gabrielclark3330 /transcribed_auspicious_president_audio Transcribed Auspicious President Audio A single-speaker English speech dataset containing 283 clips (25.36 minutes) with embedded audio and transcripts. The audio comes from Dinnerb0ner/Obama-Sample-Dataset. Transcripts were generated with Qwen/Qwen3-ASR-1.7B and may contain recognition errors. Columns audio: embedded audio bytes and source filename text: transcript source_file: original filename duration_seconds: clip duration This dataset is intended for… See the full description on the dataset page: https://huggingface.co/datasets/gabrielclark3330/transcribed_auspicious_president_audio.audioautomatic-speech-recognitionn<1K0 likes20 downloads2mo agoHugging Face05gabrielclark3330 /transcribed_auspicious_anime_girl_audio Transcribed Auspicious Anime Girl Audio A single-speaker English voice-over dataset containing 75 Arlecchino clips (18.24 minutes) with embedded audio and transcripts. Audio and transcripts were collected from the Genshin Impact Wiki Arlecchino voice-over page, revision 2123801. Only English voice-over files with nonempty transcripts are included. Columns audio: embedded audio bytes and source filename text: transcript source_file: original filename… See the full description on the dataset page: https://huggingface.co/datasets/gabrielclark3330/transcribed_auspicious_anime_girl_audio.audioautomatic-speech-recognitionn<1K0 likes19 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.