CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Shirali /ISSAI_KSC_335RS_v_1_1 Dataset Card for "ISSAI_KSC_335RS_v_1_1" Kazakh Speech Corpus (KSC) Identifier: SLR102 Summary: A crowdsourced open-source Kazakh speech corpus developed by ISSAI (330 hours) Category: Speech License: Attribution 4.0 International (CC BY 4.0) Downloads (use a mirror closer to you): ISSAI_KSC_335RS_v1.1_flac.tar.gz [19G] (speech, transcripts and metadata ) Mirrors: [US] [EU] [CN] About this resource: A crowdsourced open-source speech corpus for the Kazakh language. The KSC… See the full description on the dataset page: https://huggingface.co/datasets/Shirali/ISSAI_KSC_335RS_v_1_1.audioautomatic-speech-recognition100K<n<1M3 likes2.4k downloads4y agoHugging Face02ShiniChien /TTSDistil-Phonologygated Overview This repository contains a speech dataset developed for Text-to-Speech (TTS) research and model training. The dataset is part of an ongoing research project focused on building high-quality speech corpora for modern neural TTS systems. It is actively maintained, with continuous improvements in data quality, transcription consistency, metadata, and organization. This repository is intended to host research data used throughout the development process and is not intended… See the full description on the dataset page: https://huggingface.co/datasets/ShiniChien/TTSDistil-Phonology.audiotext-to-speech100K<n<1M1 likes858 downloads2mo agoHugging Face03ShiniChien /TTSDistil2gated Overview This repository contains a speech dataset developed for Text-to-Speech (TTS) research and model training. The dataset is part of an ongoing research project focused on building high-quality speech corpora for modern neural TTS systems. It is actively maintained, with continuous improvements in data quality, transcription consistency, metadata, and organization. This repository is intended to host research data used throughout the development process and is not intended… See the full description on the dataset page: https://huggingface.co/datasets/ShiniChien/TTSDistil2.audiotext-to-speech100K<n<1M1 likes557 downloads14d agoHugging Face04FormosanBank /ePark_tu_hua_gu_shi_pian_picture_story FormosanBank publication status This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card. FormosanBank/ePark_tu_hua_gu_shi_pian_picture_story Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum. This is a… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_tu_hua_gu_shi_pian_picture_story.audioautomatic-speech-recognition1K<n<10K0 likes225 downloads2mo agoHugging Face05Shiry /ATC_combined Dataset Card for UWB-ATCC corpus Dataset Summary The UWB-ATCC Corpus is provided provided by University of West Bohemia, Department of Cybernetics. The corpus contains recordings of communication between air traffic controllers and pilots. The speech is manually transcribed and labeled with the information about the speaker (pilot/controller, not the full identity of the person). The corpus is currently small (20 hours) but we plan to search for additional data next year.… See the full description on the dataset page: https://huggingface.co/datasets/Shiry/ATC_combined.audioautomatic-speech-recognition10K<n<100K1 likes165 downloads3y agoHugging Face06kibaraki /Shinekhen-BuryatAudio collected by Yamakoshi (Tokyo University of Foreign Studies), originally uploaded here (CC BY-SA 4.0). start_time and end_time are from the original audio clips; the audio uploaded here are already converted into per-sentence audio clips. Used in [paper] [GitHub] audioautomatic-speech-recognition1K<n<10K0 likes70 downloads1y agoHugging Face07shiprocket-ai /nirantargated Nirantar Nirantar speech dataset (22 Indian languages) in Hugging Face format. Source: AI4Bharat/Nirantar. Downloading language-wise subsets Each language is a separate configuration (subset), so you can load only one language (like ai4bharat/Rasa): from datasets import load_dataset, get_dataset_config_names # Single language (only that language's parquet is downloaded) ds = load_dataset("adjaysagar/nirantar", "hi", trust_remote_code=True) # Hindi ds["train"] #… See the full description on the dataset page: https://huggingface.co/datasets/shiprocket-ai/nirantar.audioautomatic-speech-recognition100K<n<1M3 likes35 downloads8mo agoHugging Face08ShiniChien /VoiceConvDesigngated Dataset columns Column Type Description conversation_id string Unique identifier for the conversation turn_index int32 Zero-based turn position within the conversation agent string Agent name voice string TTS voice used for this turn prompt string Full system instruction sent to Gemini Live for this speaker transcript string Text spoken in this turn duration_s float32 Approximate audio duration in seconds audio audio WAV audio (24 kHz mono 16-bit)audioautomatic-speech-recognition100K<n<1M2 likes8 downloads4mo agoHugging Face09ShiniChien /OmniEdugated Overview This dataset is being developed for training and evaluating omni-modal speech models, with a primary focus on audio-to-audio tasks. The repository is part of an ongoing research effort to build high-quality speech interaction datasets for next-generation conversational AI systems. The dataset is still under active development. Data organization, validation, annotation, and quality control are continuously being improved before a stable public release.… See the full description on the dataset page: https://huggingface.co/datasets/ShiniChien/OmniEdu.audioautomatic-speech-recognition10K<n<100K1 likes8 downloads3mo agoHugging Face10ShiniChien /OmniDistilgated Overview This dataset is being developed for training and evaluating omni-modal speech models, with a primary focus on audio-to-audio tasks. The repository is part of an ongoing research effort to build high-quality speech interaction datasets for next-generation conversational AI systems. The dataset is still under active development. Data organization, validation, annotation, and quality control are continuously being improved before a stable public release.… See the full description on the dataset page: https://huggingface.co/datasets/ShiniChien/OmniDistil.audioautomatic-speech-recognition100K<n<1M1 likes5 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.