CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Aniket-Tathe-08 /Custom_common_voice_dataset_using_RVC Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based Voice Conversion Custom common_voice_v11 corpus with a custom voice was was created using RVC(Retrieval-Based Voice Conversion) The model underwent 200 epochs of training, utilizing a total of 1 hour of audio clips. The data was scraped from Youtube. The audio in the custom generated dataset is of a YouTuber named Ajay Pandey Description license: cc0-1.0 language: - hi… See the full description on the dataset page: https://huggingface.co/datasets/Aniket-Tathe-08/Custom_common_voice_dataset_using_RVC.tabular10K<n<100K0 likes807 downloads3y agoHugging Face02shadow-wxh /VoiceCommandAudioThis is mainly used for fine tune "VoiceCommand" a speech congnition MOD dedicated for SilentHunter game series audioautomatic-speech-recognitionn<1K1 likes253 downloads2y agoHugging Face03hadou1225 /Hadou-Voice-Dataset Hadou Voice Dataset ハドウ本人が収録した、日本語音声データセットです。 このページで、特徴の異なる2種類のデータセットを公開しています。 配布データ 設定名 内容 音声数 合計時間 v1(おすすめ) Hadou Calm Voice Dataset v1。落ち着いた中音域、AIキャラクター向けボイスが多めの音声データ 966 約114.02分 v0 Hadou ITA Corpus Dataset v1。ITAコーパスを読み上げた自然な話し声 424 約38.95分 v1 には、AICAコーパス500文、ITAコーパス324文、感情・態度付き90文、同文異演技40文、強度段階12文を収録しています。 v1の詳細: v1/README.txt v0の詳細: v0/README.txt 読み込み例 from datasets import load_dataset # 新しい966音声(既定) dataset =… See the full description on the dataset page: https://huggingface.co/datasets/hadou1225/Hadou-Voice-Dataset.audiotext-to-speech1K<n<10K2 likes156 downloads1mo agoHugging Face04openpecha /tibetan-voice-benchmarkBenchmark of Tibetan Speech-To-Text dataset created by Monlam AI x Openpecha in 2024 All the transcripts have been reviewed by at least one person in addition to the original transcriber. Data was taken on 15 July 2024 02∶47∶06 PM. dept desc Count STT_AB Audio book 1000 STT_CS Children Speech 1367 STT_HS History 1000 STT_MV Tibetan Movies 1000 STT_NS Natural Speech 1000 STT_NW News 1000 STT_PC Podcast 1000 STT_TT Tibetan Teachings 1000 grade column is used to… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/tibetan-voice-benchmark.tabular1K<n<10K3 likes155 downloads8mo agoHugging Face05Arnold /hausa_common_voiceThis dataset is from the common voice corpus 7.0 using the Hausa dataset tabular1K<n<10K2 likes151 downloads5y agoHugging Face06nucklehead /ht-voice-dataset Ansanb done vwa an Kreyòl pou antrene DeepSPeech. Dataset sa a gen plis pase 7 è tan anrejistreman vwa ak prèske 100 moun an Kreyòl pou bati sistèm ASR ak TTS pou lang Kreyòl la. Pifò nan done yo soti nan "CMU Haitian Creole Speech Recognition Database" la. Done sa yo gentan filtre epi òganize pou ka antrene modèl DeepSPeech Mozilaa a. Si toutfwa ou ta bezwen jwenn plis enfòmasyon sou jan done yo ranje a epi kisa ou ka fè avèk yo, tcheke DeepSpeech Readme. audio1K<n<10K0 likes71 downloads5y agoHugging Face07fassabilf /voiceguard-competition VoiceGuard — Deepfake Audio Detection Competition Pelatnas IOAI 2026 | Task 3 of 3 Detect whether a 4-second audio clip is real human speech or AI-generated (TTS/deepfake). Submit probability scores — AUROC is the metric. Task Input: .wav audio file (4 seconds, 16 kHz mono)Output: score — probability (0–1) that the audio is fakeMetric: AUROC (Area Under ROC Curve) Dataset Split Real Fake Total Train 2,874 2,874 5,748 Test 627 627 1,254… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/voiceguard-competition.audioaudio-classification1K<n<10K0 likes71 downloads4mo agoHugging Face08etechgrid /Prompts_for_Voice_cloning_and_TTStext100K<n<1M7 likes70 downloads3y agoHugging Face09mbazaNLP /common-voice-kinyarwanda-english-dataset Kinyarwanda-English Commonvoice dataset A compilation of Kinyarwanda-english dataset to be used to train multi-lingual ASR Note: The audio dataset shall be added in the future text100K<n<1M0 likes58 downloads4y agoHugging Face10iBlessi /home-assistant-local-llm-voice-benchmark Home Assistant Local LLM Voice Benchmark Per-model tool-call accuracy and component latency for running a Home Assistant voice assistant against local LLMs. Measured per-model tool-call accuracy and component latency for running a Home Assistant voice assistant against local LLMs. Broken out by pipeline component rather than reported as one opaque round trip, so you can tell whether your latency is wake-word, speech-to-text, the model, or text-to-speech before you go optimizing… See the full description on the dataset page: https://huggingface.co/datasets/iBlessi/home-assistant-local-llm-voice-benchmark.tabularn<1K0 likes47 downloads25d agoHugging Face11Voicemod /librispeech40tabular1K<n<10K3 likes43 downloads4y agoHugging Face12ShiroOnigami23 /emotion-voice-dataset emotion voice dataset Developed by Aryan Singh Chandel (Shiro) at Rustamji Institute of Technology (RJIT). 📝 Overview This repository contains assets for emotion voice dataset. It is a professional research component of the Shiro AI ecosystem. 🚀 Status The core files are live. Detailed usage instructions and technical benchmarks are currently being compiled for the elite release. audio1K<n<10K5 likes40 downloads9mo agoHugging Face13alex73 /mozilla-common-voice-23-bel-texts-exporttabulartext-generation100K<n<1M0 likes39 downloads10mo agoHugging Face14ZamAI-Pashto /zamai-pashto-voice2voice ZamAI Pashto Voice2Voice This dataset contains Pashto voice-to-voice preparation metadata for speech and translation experiments. It focuses on Pashto speech records, dialect information, transcript text, and a small viewer-ready sample manifest. Configs from datasets import load_dataset metadata = load_dataset("ZamAI-Pashto/zamai-pashto-voice2voice", "metadata") sample = load_dataset("ZamAI-Pashto/zamai-pashto-voice2voice", "viewer_sample") Files… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-voice2voice.tabularautomatic-speech-recognitionn<1K0 likes37 downloads2mo agoHugging Face15sdialog /voices-kokorotextn<1K0 likes34 downloads11mo agoHugging Face16daroArt /col_voiceaudio1K<n<10K0 likes28 downloads2y agoHugging Face17MonlamAI /Bo-voice-v1.0.0 Tibetan STT Benchmark Model Card Bo-voice-v1.0.0 is a high-fidelity benchmark for Tibetan Speech-to-Text (STT) technology. It provides a rigorous, multi-domain evaluation set to measure Automatic Speech Recognition (ASR) performance across diverse acoustic environments and speaking styles. ### Dataset Overview Snapshot Date: 15 July 2024, 02:47:06 PM Total Samples: 8,367 audio-transcript pairs. Verification: Every transcript has been reviewed by at least one expert in… See the full description on the dataset page: https://huggingface.co/datasets/MonlamAI/Bo-voice-v1.0.0.tabular1K<n<10K0 likes27 downloads5mo agoHugging Face18tasal9 /zamai-pashto-voice2voice ZamAI Pashto Voice2Voice Languages: psLicense: cc-by-4.0Task categories: automatic-speech-recognition, audio-to-audioSize categories: n<1K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for automatic-speech-recognition, audio-to-audio tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/zamai-pashto-voice2voice") print(dataset) Configs default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/zamai-pashto-voice2voice.audioautomatic-speech-recognition1K<n<10K1 likes27 downloads2mo agoHugging Face19voice-bench-submission /voxclinbench VoxClinBench (Hugging Face dataset mirror) Cross-lingual, cross-disease clinical voice biomarker benchmark. This Hugging Face dataset mirror ships the split manifests, Croissant metadata, datasheet, and reference baseline prediction CSVs. Raw audio is NOT distributed here. Each of the five upstream corpora must be obtained directly from its provider under that corpus's data use agreement (DUA). See the GitHub mirror for the evaluation harness and the voxbench fetch CLI.… See the full description on the dataset page: https://huggingface.co/datasets/voice-bench-submission/voxclinbench.textaudio-classification1K<n<10K0 likes24 downloads5mo agoHugging Face20CocoaRain /common_voice_13_0_zh_pseudo_labelledtext1K<n<10K0 likes22 downloads3y agoHugging Face21CS-224s /common-voicetabular10K<n<100K0 likes22 downloads2y agoHugging Face22Lkhagvasurenam /common_voice_13_0_mn_pseudo_test_smalltextn<1K0 likes21 downloads3y agoHugging Face23laion /voice-acting-instructionstext100K<n<1M1 likes21 downloads10mo agoHugging Face24javadr /common_voice_16_0_fa_pseudo_labelledtext10K<n<100K0 likes20 downloads3y agoHugging Face25bjak /common_voice_13_0_thai_small_pseudo_labelledtextn<1K0 likes19 downloads3y agoHugging Face26ro-anderson /linkedin-top-voices-market-alpaca linkedin-top-voices-market-alpaca Dataset Dataset containing 300 records for fine-tuning language models. Columns instruction input output input_tokens output_tokens input_cost output_cost total_cost tabularn<1K1 likes17 downloads1y agoHugging Face27tor24 /smart-home-voice-commands-v1 Smart Home Voice Commands V1 Description Dataset of smart home voice assistant commands labeled by device control category. Data Fields text: Voice command intent: Device action label Intents lights_on lights_off increase_temperature decrease_temperature play_music License CC-BY-4.0 textn<1K0 likes17 downloads7mo agoHugging Face28xiaochen233 /mdd-voice-biomarker-data mdd-voice-biomarker-data Small task-specific mirror of public input files used by Terminal Bench Science task mdd-voice-biomarker. The files are redistributed here to make benchmark Docker builds faster and more reproducible. See task instructions and original upstream sources for dataset-specific citation and licensing details. Contents data/: input files downloaded by the task Dockerfiles. SHA256SUMS.txt: SHA-256 checksums for all files under data/. tabularn<1K0 likes16 downloads3mo agoHugging Face29omarsou /common_voice_16_1_spanish_test_set Dataset Card for Common Voice Corpus 16 Spanish Dataset Acknowledgement The dataset belongs to COMMON VOICE MOZILLA FOUNDATION. I just uploaded the spanish test set (from HERE : https://huggingface.co/datasets/mozilla-foundation/common_voice_16_1/tree/main) Dataset Summary The Common Voice dataset consists of a unique MP3 and corresponding text file. Languages Spanish How to use The datasets library allows you to load and pre-process… See the full description on the dataset page: https://huggingface.co/datasets/omarsou/common_voice_16_1_spanish_test_set.tabular10K<n<100K1 likes15 downloads2y agoHugging Face30Mike136 /common_voice_16_1_hi_pseudo_labelledtext1K<n<10K0 likes11 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.