CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01EmilyNguyen235 /data-voice-vietnamese-restaurant-quan-oc Vietnamese Restaurant Order Speech This dataset contains Vietnamese spoken restaurant orders paired with text transcripts. Each utterance typically includes a table number, item quantities, dishes, drinks, and add-ons. Dataset Structure Files are split into subdirectories by filename-derived speaker_code to satisfy Hugging Face repository file-count limits: metadata.csv: one row per audio sample. audio/{speaker_code}/*.wav: mono WAV audio files.… See the full description on the dataset page: https://huggingface.co/datasets/EmilyNguyen235/data-voice-vietnamese-restaurant-quan-oc.audioautomatic-speech-recognition1K<n<10K1 likes325 downloads3mo agoHugging Face02instinct-org /yt4_chunked_speech_restorised_tts_traingated yt4_chunked_speech_restorised_tts_train This is a gated Russian TTS training dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_chunked_speech_restorised_tts_train.audiotext-to-speech0 likes248 downloads4mo agoHugging Face03instinct-org /yt_chunked_speech_restorised_tts_traingated yt_chunked_speech_restorised_tts_train This is a gated Russian TTS training dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt_chunked_speech_restorised_tts_train.audiotext-to-speech0 likes206 downloads4mo agoHugging Face04instinct-org /yt3_chunked_speech_restorised_tts_traingated yt3_chunked_speech_restorised_tts_train This is a gated Russian TTS training dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Prepared for TTS… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_speech_restorised_tts_train.audiotext-to-speech0 likes198 downloads4mo agoHugging Face05instinct-org /yt1_chunked_speech_restorised_tts_traingated yt1_chunked_speech_restorised_tts_train This is a gated Russian TTS training dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Prepared for TTS… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt1_chunked_speech_restorised_tts_train.audiotext-to-speech0 likes146 downloads4mo agoHugging Face06instinct-org /yt2_chunked_speech_restorised_tts_traingated yt2_chunked_speech_restorised_tts_train This is a gated Russian TTS training dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Prepared for TTS… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_speech_restorised_tts_train.audiotext-to-speech0 likes138 downloads4mo agoHugging Face07instinct-org /miscellaneous_yt_chunked_speech_restorised_tts_traingated miscellaneous_yt_chunked_speech_restorised_tts_train This is a gated Uzbek TTS training dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Prepared for… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_speech_restorised_tts_train.audiotext-to-speech0 likes106 downloads4mo agoHugging Face08KaniTTS-research-team /dataset-from-restoreaudio1K<n<10K0 likes82 downloads2mo agoHugging Face09archivartaunik /cv-corpus21_be-sidon-restored-1000 cv-corpus21_be-sidon-restored-1000 Прыклад датасэта з 639 запісамі (Belarusian, Common Voice validated), дзе audio — адноўлены WAV (48 кГц), original_audio — арыгінальны кліп, а таксама sentence і speaker. Створана: 2025-09-26. audion<1K0 likes20 downloads1y agoHugging Face10yagmurx /ataturk_voice_no_restorationaudion<1K1 likes19 downloads3y agoHugging Face11figfig /restaurant_order_HSR_test Dataset Card for "restaurant_order_HSR_test" More Information needed audion<1K0 likes13 downloads4y agoHugging Face12instinct-org /tbp_chunked_speech_restorisedgated tbp_chunked_speech_restorised This is a gated Russian speech-restorised chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: ru (Russian) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked_speech_restorised.audioautomatic-speech-recognition100K<n<1M0 likes13 downloads4mo agoHugging Face13RHEZLOUNE /darija-restaurant-audioaudion<1K0 likes11 downloads3mo agoHugging Face14archivartaunik /be-sidon-restored-sample-10-fixed be-sidon-restored-sample-10-fixed Прыклад датасэта з 10 запісамі (Belarusian, Common Voice validated), дзе audio — адноўлены WAV (48 кГц), original_audio — арыгінальны кліп, а таксама sentence і speaker. Створана: 2025-09-25. audion<1K0 likes9 downloads1y agoHugging Face15infinite-learning-station /whisper_restaurant_trainingaudion<1K0 likes8 downloads9mo agoHugging Face16instinct-org /espeech_podcasts_chunked_speech_restorised_tts_train_clone_pairsgated espeech_podcasts_chunked_speech_restorised_tts_train_clone_pairs This is a gated Russian TTS training clone-pair dataset. It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows. Language Primary language: ru (Russian) Contents audios/shard-*.tar: tokenized audio shards txts/shard-*.jsonl: per-example metadata and text fields data.lst: repository-relative shard manifest… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_speech_restorised_tts_train_clone_pairs.audiotext-to-speech0 likes8 downloads4mo agoHugging Face17instinct-org /default_voices_chunked_speech_restorisedgated default_voices_chunked_speech_restorised This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_speech_restorised.audioautomatic-speech-recognition100K<n<1M0 likes6 downloads4mo agoHugging Face18instinct-org /cv_chunked_speech_restorisedgated cv_chunked_speech_restorised This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_speech_restorised.audioautomatic-speech-recognition10K<n<100K0 likes6 downloads1mo agoHugging Face19instinct-org /default_voices_chunked_speech_restorised_tts_train_clone_pairsgated default_voices_chunked_speech_restorised_tts_train_clone_pairs This is a gated Uzbek TTS training clone-pair dataset. It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows. Language Primary language: uz (Uzbek) Contents audios/shard-*.tar: tokenized audio shards txts/shard-*.jsonl: per-example metadata and text fields data.lst: repository-relative shard manifest tokenized_dataset_summary.json:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_speech_restorised_tts_train_clone_pairs.audiotext-to-speech0 likes6 downloads4mo agoHugging Face20instinct-org /yt2_chunked_speech_restorised_tts_train_clone_pairsgated yt2_chunked_speech_restorised_tts_train_clone_pairs This is a gated Russian TTS training clone-pair dataset. It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows. Language Primary language: ru (Russian) Contents audios/shard-*.tar: tokenized audio shards txts/shard-*.jsonl: per-example metadata and text fields data.lst: repository-relative shard manifest tokenized_dataset_summary.json: upload-time… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_speech_restorised_tts_train_clone_pairs.audiotext-to-speech0 likes6 downloads4mo agoHugging Face21instinct-org /yt3_chunked_speech_restorised_tts_train_clone_pairsgated yt3_chunked_speech_restorised_tts_train_clone_pairs This is a gated Russian TTS training clone-pair dataset. It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows. Language Primary language: ru (Russian) Contents audios/shard-*.tar: tokenized audio shards txts/shard-*.jsonl: per-example metadata and text fields data.lst: repository-relative shard manifest tokenized_dataset_summary.json: upload-time… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_speech_restorised_tts_train_clone_pairs.audiotext-to-speech0 likes6 downloads4mo agoHugging Face22figfig /restaurant_order_local_test_colabgated Dataset Card for "restaurant_order_local_test_colab" More Information needed audion<1K0 likes5 downloads4y agoHugging Face23yagmurx /ataturk_voice_restoratedaudion<1K2 likes5 downloads3y agoHugging Face24instinct-org /audiobook_chunked_speech_restorisedgated audiobook_chunked_speech_restorised This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_speech_restorised.audioautomatic-speech-recognition1M<n<10M0 likes5 downloads4mo agoHugging Face25instinct-org /tbp_chunked_speech_restorised_tts_train_clone_pairsgated tbp_chunked_speech_restorised_tts_train_clone_pairs This is a gated Russian TTS training clone-pair dataset. It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows. Language Primary language: ru (Russian) Contents audios/shard-*.tar: tokenized audio shards txts/shard-*.jsonl: per-example metadata and text fields data.lst: repository-relative shard manifest tokenized_dataset_summary.json: upload-time… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked_speech_restorised_tts_train_clone_pairs.audiotext-to-speech0 likes5 downloads4mo agoHugging Face26instinct-org /audiobook_chunked_speech_restorised_tts_train_clone_pairsgated audiobook_chunked_speech_restorised_tts_train_clone_pairs This is a gated Uzbek TTS training clone-pair dataset. It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows. Language Primary language: uz (Uzbek) Contents audios/shard-*.tar: tokenized audio shards txts/shard-*.jsonl: per-example metadata and text fields data.lst: repository-relative shard manifest tokenized_dataset_summary.json:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_speech_restorised_tts_train_clone_pairs.audiotext-to-speech0 likes5 downloads4mo agoHugging Face27instinct-org /miscellaneous_yt_chunked_speech_restorised_tts_train_clone_pairsgated miscellaneous_yt_chunked_speech_restorised_tts_train_clone_pairs This is a gated Uzbek TTS training clone-pair dataset. It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows. Language Primary language: uz (Uzbek) Contents audios/shard-*.tar: tokenized audio shards txts/shard-*.jsonl: per-example metadata and text fields data.lst: repository-relative shard manifest tokenized_dataset_summary.json:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_speech_restorised_tts_train_clone_pairs.audiotext-to-speech0 likes5 downloads4mo agoHugging Face28instinct-org /yt4_chunked_speech_restorised_tts_train_clone_pairsgated yt4_chunked_speech_restorised_tts_train_clone_pairs This is a gated Russian TTS training clone-pair dataset. It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows. Language Primary language: ru (Russian) Contents audios/shard-*.tar: tokenized audio shards txts/shard-*.jsonl: per-example metadata and text fields data.lst: repository-relative shard manifest tokenized_dataset_summary.json: upload-time… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_chunked_speech_restorised_tts_train_clone_pairs.audiotext-to-speech0 likes5 downloads4mo agoHugging Face29archivartaunik /cv-corpus20_be-sidon-restored-1000 cv-corpus20_be-sidon-restored-1000 Прыклад датасэта з 1000 запісамі (Belarusian, Common Voice validated), дзе audio — адноўлены WAV (48 кГц), original_audio — арыгінальны кліп, а таксама sentence і speaker. Створана: 2025-09-26. audio1K<n<10K0 likes4 downloads1y agoHugging Face30infinite-learning-station /whisper_indian_restaurant_trainingaudio10K<n<100K0 likes4 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.