CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Mozilla /standard_chattextn<1K0 likes566 downloads29d agoHugging Face02Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_3 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_3" More Information needed audio10K<n<100K0 likes486 downloads3y agoHugging Face03RidheshBhati /Indic_Mozilla_TTS Indic TTS Dataset Hub (Mozilla) Validated audio–text pairs from Mozilla Common Voice for multiple Indic languages (and English). Select the language from the Subset dropdown in the Dataset Viewer. Columns audio: WAV audio clip (16 kHz, embedded bytes) text: TTS-ready transcription duration: audio length in seconds speaking_rate: characters per second audio100K<n<1M0 likes329 downloads7mo agoHugging Face04Mozilla /standard_chat_tool_calling_generaltextn<1K1 likes326 downloads4mo agoHugging Face05Mozilla /link_tab_hallucination_eval link_tab_hallucination_eval Curated eval for Firefox AI Window link-hallucination and tab-read failure patterns (false_login, needless_fetch, describe_without_reading), plus link-hallucination prompts. Tab-read cases are pre-seeded 2-turn threads: a get_page_content tool-call + its result (a frozen page snapshot) are baked into the message thread so predictions are reproducible (no live fetch), while the final scorable user turn still shows the real tab URL. 139 rows; fields:… See the full description on the dataset page: https://huggingface.co/datasets/Mozilla/link_tab_hallucination_eval.textn<1K0 likes304 downloads2mo agoHugging Face06Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_2 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_2" More Information needed audio10K<n<100K0 likes288 downloads3y agoHugging Face07Mozilla /standard_chat_manage_tabs_adversarialtextn<1K0 likes274 downloads2mo agoHugging Face08Mozilla /standard_chat_longconv_languagetextn<1K0 likes259 downloads2mo agoHugging Face09EYEDOL /mozilla_commonvoice_naijaHausa1_preprocessed_train_batch_1audio10K<n<100K0 likes234 downloads1y agoHugging Face10Mozilla /flickr30k-transformed-captionsThis is a "de-biased" version of https://huggingface.co/datasets/nlphuji/flickr30k dataset. We've added a few extra columns: alt_text: the captions rewritten by calling the meta-llama/Meta-Llama-3-8B-Instruct LLM grade: a measure of redability using the readability library Learn more about why and how we did it here : https://github.com/mozilla/distilvit/blob/main/docs/fighting_bias.md See the code here : https://github.com/mozilla/distilvit/blob/main/distilvit/curate.py For the licence… See the full description on the dataset page: https://huggingface.co/datasets/Mozilla/flickr30k-transformed-captions.image10K<n<100K8 likes213 downloads2y agoHugging Face11EYEDOL /mozilla_commonvoice_Swahili_preprocessed_train_batch_1audio10K<n<100K0 likes204 downloads1y agoHugging Face12EYEDOL /mozilla_commonvoice_Swahili_preprocessed_train_batch_3audio10K<n<100K0 likes195 downloads1y agoHugging Face13yakhyo /mozilla-common-voice-uzbek 🗣️ Mozilla Common Voice (Uzbek) — Cleaned & Normalized This dataset is a refined version of mozilla-foundation/common_voice_17_0, containing only Uzbek language voice recordings, and enriched with preprocessing steps for better usability in training ASR models. 🔍 Dataset Overview This version focuses exclusively on Uzbek audio samples and includes the following modifications: 🎯 Filtered to include only Uzbek examples. ✨ Normalized text field added under the key text.… See the full description on the dataset page: https://huggingface.co/datasets/yakhyo/mozilla-common-voice-uzbek.audio100K<n<1M4 likes192 downloads1y agoHugging Face14Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_5 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_5" More Information needed audio10K<n<100K0 likes181 downloads3y agoHugging Face15Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_1 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_1" More Information needed audio10K<n<100K0 likes160 downloads3y agoHugging Face16EYEDOL /mozilla_commonvoice_Swahili_preprocessed_train_batch_4audio10K<n<100K0 likes157 downloads1y agoHugging Face17Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_6 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_6" More Information needed audio10K<n<100K0 likes152 downloads3y agoHugging Face18Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_4 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_4" More Information needed audio10K<n<100K0 likes146 downloads3y agoHugging Face19EYEDOL /mozilla_commonvoice_Swahili_preprocessed_train_batch_2audio10K<n<100K0 likes136 downloads1y agoHugging Face20EYEDOL /mozilla_commonvoice_Swahili_preprocessed_train_batch_5audio1K<n<10K0 likes110 downloads1y agoHugging Face21Mozilla /alt-text-validationThis dataset contains images and alt text from various sources. It is used to control the quality of https://huggingface.co/Mozilla/distilvit using the https://github.com/mozilla/checkvite application This application let users try out the model on the images and classify them. The dataset is then updated. When an image is marked as need_training it will be use to fine-tune the model to fix some of its inaccuracies image1K<n<10K7 likes104 downloads2y agoHugging Face22EYEDOL /mozilla_commonvoice_Arabic_preprocessed_train_batch_2audio1K<n<10K0 likes102 downloads1y agoHugging Face23Mozilla /flickr30k-transformed-captions-gpt4oThis is a "de-biased" version of https://huggingface.co/datasets/nlphuji/flickr30k dataset. The new alt_text column was produced by GPT-4o using the following script : https://github.com/mozilla/distilvit/blob/main/distilvit/curate_gpt.py Learn more about why and how we did it here : https://github.com/mozilla/distilvit/blob/main/docs/fighting_bias.md See the code here : https://github.com/mozilla/distilvit/blob/main/distilvit/curate_gpt.py For the licence, see the original dataset. image10K<n<100K3 likes97 downloads2y agoHugging Face24Mozilla /smart-form-fill-value-generationtextn<1K0 likes97 downloads29d agoHugging Face25EYEDOL /mozilla_commonvoice_Swahili_preprocessed_train_batch_6audio1K<n<10K0 likes96 downloads1y agoHugging Face26Mozilla /smart-form-fill-e2etextn<1K0 likes86 downloads29d agoHugging Face27fibleep /flemish-mozilla-common-voiceA port of dutch-vl-tts [https://github.com/r-dh/dutch-vl-tts] to hugging face. Uses 15 000 samples of a male Dutch Flemish voice. Extracted from Mozilla Common Voice project [https://github.com/common-voice/common-voice/tree/master/server/data/nl]. audio10K<n<100K1 likes81 downloads2y agoHugging Face28EYEDOL /mozilla_commonvoice_Arabic_preprocessed_train_batch_1audio1K<n<10K1 likes74 downloads1y agoHugging Face29Mozilla /pexels-gpt4oImages collected from Pexels, using 1000 images following 5 categories: nudes war group animals smoking See https://www.pexels.com/license/ for the license They were then annotated using gpt4-o, see https://github.com/mozilla/distilvit/blob/main/distilvit/gpt4.py image1K<n<10K4 likes70 downloads2y agoHugging Face30MatheusMarquesEiras /mozilla-common-voice-converted-to-parquet-pttabular10K<n<100K0 likes64 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.