CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01artur-muratov /multilingual-speech-commands-15lang Multilingual Speech Commands Dataset (15 Languages, Augmented) This dataset contains augmented speech command samples in 15 languages, derived from multiple public datasets. Only commands that overlap with the Google Speech Commands (GSC) vocabulary are included, making the dataset suitable for multilingual keyword spotting tasks aligned with GSC-style classification. Audio samples have been augmented using standard audio techniques to improve model robustness (e.g., time-shifting… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-15lang.audio1M<n<10M16 likes80k downloads1y agoHugging Face02artur-muratov /multilingual-speech-commands-3lang-raw Multilingual Speech Commands Dataset (3 Languages, Raw) This dataset is a curated subset of previously published speech command datasets in Kazakh, Tatar, and Russian. It is intended for use in multilingual speech command recognition and keyword spotting tasks. No data augmentation has been applied. All files are included in their original form as released in the cited works below. This repository simply reorganizes them for convenience and accessibility. Languages… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-3lang-raw.audio1K<n<10K1 likes3.9k downloads1y agoHugging Face03mazkooleg /0-9up_google_speech_commands_augmented_raw Dataset Card for "google_speech_commands_augmented_raw_fixed" More Information needed audio1M<n<10M0 likes3k downloads4y agoHugging Face04Codec-SUPERB /fluent_speech_commands_synth Dataset Card for "fluent_speech_commands_synth" More Information needed audio100K<n<1M1 likes1.3k downloads3y agoHugging Face05pollen-robotics /speech-commands-v0.02 Speech Commands Dataset v0.02 This is a re-hosted copy of the Google Speech Commands v0.02 dataset in Parquet format for compatibility with the Hugging Face Dataset Viewer. ⚠️ Credits This dataset was created by Pete Warden / Google. All credit goes to the original authors and the crowdsourcing contributors. Original source: http://download.tensorflow.org/data/speech_commands_v0.02.tar.gz Paper: Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/speech-commands-v0.02.audioaudio-classification100K<n<1M0 likes1.1k downloads3mo agoHugging Face06soerenray /speech_commands_enriched_and_annotated Dataset Summary 📊 Data-centric AI principles have become increasingly important for real-world use cases.At Renumics we believe that classical benchmark datasets and competitions should be extended to reflect this development. 🔍 This is why we are publishing benchmark datasets with application-specific enrichments (e.g. embeddings, baseline results, uncertainties, label error scores). We hope this helps the ML community in the following ways: Enable new researchers to quickly… See the full description on the dataset page: https://huggingface.co/datasets/soerenray/speech_commands_enriched_and_annotated.audio10K<n<100K2 likes569 downloads3y agoHugging Face07CodecSR /fluent_speech_commands_femaleaudio10K<n<100K1 likes436 downloads2y agoHugging Face08hoangbang /hey-computer-speech-commands Hey Computer: Speech Command Recognition Dataset Summary A public, viewer-ready educational challenge dataset. Host-only scoring data and hidden targets are excluded. Splits Split Examples Description train 13,192 Labeled training data test 3,295 Public inputs with withheld target labels or annotations Data Fields Field Type audio Audio id string label string (test sentinel: unlabeled)… See the full description on the dataset page: https://huggingface.co/datasets/hoangbang/hey-computer-speech-commands.audioaudio-classification10K<n<100K0 likes378 downloads2mo agoHugging Face09AigizK /bashkort_commands_omnivoice Bashkort Commands OmniVoice Partial eleven-label command snapshot generated with k2-fsa/OmniVoice using the same cross-lingual voice-cloning recipe as AigizK/homai_wake_word_omnivoice. Generation was stopped at the user's request after 41,525 complete reference groups had been committed. For every included reference row from the train split of: bond005/sova_rudevices the dataset contains one recording of every command: Айвика — Russian Айвикә — Bashkir Айһылыу — Bashkir… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/bashkort_commands_omnivoice.audioaudio-classification100K<n<1M0 likes308 downloads2mo agoHugging Face10danjacobellis /speech_commands_v2audio100K<n<1M0 likes274 downloads2y agoHugging Face11CodecSR /fluent_speech_commands_maleaudio10K<n<100K0 likes273 downloads2y agoHugging Face12artur-muratov /multilingual-speech-commands-15lang-zip Multilingual Speech Commands Dataset (15 Languages, Augmented) This dataset contains augmented speech command samples in 15 languages, derived from multiple public datasets. Only commands that overlap with the Google Speech Commands (GSC) vocabulary are included, making the dataset suitable for multilingual keyword spotting tasks aligned with GSC-style classification. Audio samples have been augmented using standard audio techniques to improve model robustness (e.g., time-shifting… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-15lang-zip.audio1M<n<10M1 likes271 downloads1y agoHugging Face13renumics /speech_commands_enrichedThis is a set of one-second .wav audio files, each containing a single spoken English word or background noise. These words are from a small set of commands, and are spoken by a variety of different speakers. This data set is designed to help train simple machine learning models. This dataset is covered in more detail at [https://arxiv.org/abs/1804.03209](https://arxiv.org/abs/1804.03209). Version 0.01 of the data set (configuration `"v0.01"`) was released on August 3rd 2017 and contains 64,727 audio files. In version 0.01 thirty different words were recoded: "Yes", "No", "Up", "Down", "Left", "Right", "On", "Off", "Stop", "Go", "Zero", "One", "Two", "Three", "Four", "Five", "Six", "Seven", "Eight", "Nine", "Bed", "Bird", "Cat", "Dog", "Happy", "House", "Marvin", "Sheila", "Tree", "Wow". In version 0.02 more words were added: "Backward", "Forward", "Follow", "Learn", "Visual". In both versions, ten of them are used as commands by convention: "Yes", "No", "Up", "Down", "Left", "Right", "On", "Off", "Stop", "Go". Other words are considered to be auxiliary (in current implementation it is marked by `True` value of `"is_unknown"` feature). Their function is to teach a model to distinguish core words from unrecognized ones. This version is not yet supported. The `_silence_` class contains a set of longer audio clips that are either recordings or a mathematical simulation of noise.audioaudio-classification10K<n<100K3 likes254 downloads3y agoHugging Face14negfir /speech_commands_pitchaudio10K<n<100K0 likes241 downloads10mo agoHugging Face15renumics /speech_commands_enrichment_only Dataset Card for SpeechCommands Dataset Summary 📊 Data-centric AI principles have become increasingly important for real-world use cases.At Renumics we believe that classical benchmark datasets and competitions should be extended to reflect this development. 🔍 This is why we are publishing benchmark datasets with application-specific enrichments (e.g. embeddings, baseline results, uncertainties, label error scores). We hope this helps the ML community in the… See the full description on the dataset page: https://huggingface.co/datasets/renumics/speech_commands_enrichment_only.audioaudio-classification100K<n<1M0 likes233 downloads3y agoHugging Face16Hunzla /simplified_google_speech_commands_wav2vec2_960haudio10K<n<100K1 likes222 downloads3y agoHugging Face17lugan /Syntts-Commands-Media-Dataset SynTTS-Commands: A Multilingual Synthetic Speech Command Dataset 📖 Introduction SynTTS-Commands is a large-scale, multilingual synthetic speech command dataset specifically designed for low-power Keyword Spotting (KWS) and speech command recognition tasks. As presented in the paper SynTTS-Commands: A Public Dataset for On-Device KWS via TTS-Synthesized Multilingual Speech, this dataset is generated using advanced Text-to-Speech (TTS) technologies, aiming to… See the full description on the dataset page: https://huggingface.co/datasets/lugan/Syntts-Commands-Media-Dataset.audioaudio-classification1 likes219 downloads8mo agoHugging Face18Abdelkareem /arabic_commands_detection Dataset Card for "arabic_commands_detection" More Information needed audio10K<n<100K0 likes193 downloads3y agoHugging Face19WhissleAI /synthetic_speech_commands_PA_taggedaudioautomatic-speech-recognition1K<n<10K0 likes178 downloads1y agoHugging Face20Codec-SUPERB /fluent_speech_commands_test_subset_synth Dataset Card for "fluent_speech_commands_test_subset_synth" More Information needed audio10K<n<100K0 likes174 downloads3y agoHugging Face21Angeriod /in_car_commands_26 Dataset Card for "in_car_commands_26" More Information needed audio10K<n<100K0 likes173 downloads2y agoHugging Face22tomas-gajarsky /speech-commands-lt Speech Commands-LT (Long Tail) Long-tail variants of Google Speech Commands v0.02 for benchmarking imbalanced audio classification. Dataset Summary Speech Commands v0.02 (Warden, 2018) contains 35 spoken word classes with ~1,200-3,200 samples each. This dataset applies exponential decay to the training set to create long-tail distributions with varying imbalance ratios, simulating real-world class imbalance in audio classification. The _silence_ class (label 35) is… See the full description on the dataset page: https://huggingface.co/datasets/tomas-gajarsky/speech-commands-lt.audioaudio-classification100K<n<1M0 likes116 downloads6mo agoHugging Face23beeneptune /speech_commands--> This is an exact copy of google/speech_commands adapted to be usable with recent datasets 🤗 versions (no remote code). <-- Dataset Card for SpeechCommands Dataset Summary This is a set of one-second .wav audio files, each containing a single spoken English word or background noise. These words are from a small set of commands, and are spoken by a variety of different speakers. This data set is designed to help train simple machine learning models. It is covered in… See the full description on the dataset page: https://huggingface.co/datasets/beeneptune/speech_commands.audioaudio-classification10K<n<100K0 likes115 downloads7mo agoHugging Face24Hunzla /google-speech-commands-wav2vec2-960haudio10K<n<100K1 likes107 downloads3y agoHugging Face25luvox-ai /commands_vi_voice_syntheticgatedaudio10K<n<100K0 likes79 downloads7mo agoHugging Face26Hunzla /simplified-google-speech-commands-wav2vec2-960haudio10K<n<100K0 likes72 downloads3y agoHugging Face27negfir /speech_commands_augmented_alpha200audio100K<n<1M0 likes62 downloads10mo agoHugging Face28jacobstgerm /speech_commands_10labels_50000synthaudio10K<n<100K0 likes59 downloads9mo agoHugging Face29Angeriod /in_car_commands_60 Dataset Card for "in_car_commands_60" More Information needed audio10K<n<100K0 likes56 downloads2y agoHugging Face30mteb /speech-commands-miniaudio10K<n<100K0 likes54 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.