datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
standard_chatmozilla_commonvoice_hackathon_preprocessed_train_batch_3
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_3"
More Information needed
Indic_Mozilla_TTS
Indic TTS Dataset Hub (Mozilla)
Validated audio–text pairs from Mozilla Common Voice for multiple Indic languages (and English).
Select the language from the Subset dropdown in the Dataset Viewer.
Columns
audio: WAV audio clip (16 kHz, embedded bytes)
text: TTS-ready transcription
duration: audio length in seconds
speaking_rate: characters per second
standard_chat_tool_calling_generallink_tab_hallucination_eval
link_tab_hallucination_eval
Curated eval for Firefox AI Window link-hallucination and tab-read failure patterns
(false_login, needless_fetch, describe_without_reading), plus link-hallucination prompts.
Tab-read cases are pre-seeded 2-turn threads: a get_page_content tool-call + its result
(a frozen page snapshot) are baked into the message thread so predictions are reproducible
(no live fetch), while the final scorable user turn still shows the real tab URL.
139 rows; fields:… See the full description on the dataset page: https://huggingface.co/datasets/Mozilla/link_tab_hallucination_eval.mozilla_commonvoice_hackathon_preprocessed_train_batch_2
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_2"
More Information needed
standard_chat_manage_tabs_adversarialstandard_chat_longconv_languagemozilla_commonvoice_naijaHausa1_preprocessed_train_batch_1flickr30k-transformed-captionsThis is a "de-biased" version of https://huggingface.co/datasets/nlphuji/flickr30k dataset. We've added a few extra columns:
alt_text: the captions rewritten by calling the meta-llama/Meta-Llama-3-8B-Instruct LLM
grade: a measure of redability using the readability library
Learn more about why and how we did it here : https://github.com/mozilla/distilvit/blob/main/docs/fighting_bias.md
See the code here : https://github.com/mozilla/distilvit/blob/main/distilvit/curate.py
For the licence… See the full description on the dataset page: https://huggingface.co/datasets/Mozilla/flickr30k-transformed-captions.mozilla_commonvoice_Swahili_preprocessed_train_batch_1mozilla_commonvoice_Swahili_preprocessed_train_batch_3mozilla-common-voice-uzbek
🗣️ Mozilla Common Voice (Uzbek) — Cleaned & Normalized
This dataset is a refined version of mozilla-foundation/common_voice_17_0, containing only Uzbek language voice recordings, and enriched with preprocessing steps for better usability in training ASR models.
🔍 Dataset Overview
This version focuses exclusively on Uzbek audio samples and includes the following modifications:
🎯 Filtered to include only Uzbek examples.
✨ Normalized text field added under the key text.… See the full description on the dataset page: https://huggingface.co/datasets/yakhyo/mozilla-common-voice-uzbek.mozilla_commonvoice_hackathon_preprocessed_train_batch_5
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_5"
More Information needed
mozilla_commonvoice_hackathon_preprocessed_train_batch_1
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_1"
More Information needed
mozilla_commonvoice_Swahili_preprocessed_train_batch_4mozilla_commonvoice_hackathon_preprocessed_train_batch_6
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_6"
More Information needed
mozilla_commonvoice_hackathon_preprocessed_train_batch_4
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_4"
More Information needed
mozilla_commonvoice_Swahili_preprocessed_train_batch_2mozilla_commonvoice_Swahili_preprocessed_train_batch_5alt-text-validationThis dataset contains images and alt text from various sources.
It is used to control the quality of https://huggingface.co/Mozilla/distilvit using the https://github.com/mozilla/checkvite application
This application let users try out the model on the images and classify them. The dataset is then updated.
When an image is marked as need_training it will be use to fine-tune the model to fix some of its inaccuracies
mozilla_commonvoice_Arabic_preprocessed_train_batch_2flickr30k-transformed-captions-gpt4oThis is a "de-biased" version of https://huggingface.co/datasets/nlphuji/flickr30k dataset.
The new alt_text column was produced by GPT-4o using the following script : https://github.com/mozilla/distilvit/blob/main/distilvit/curate_gpt.py
Learn more about why and how we did it here : https://github.com/mozilla/distilvit/blob/main/docs/fighting_bias.md
See the code here : https://github.com/mozilla/distilvit/blob/main/distilvit/curate_gpt.py
For the licence, see the original dataset.
smart-form-fill-value-generationmozilla_commonvoice_Swahili_preprocessed_train_batch_6smart-form-fill-e2eflemish-mozilla-common-voiceA port of dutch-vl-tts [https://github.com/r-dh/dutch-vl-tts] to hugging face.
Uses 15 000 samples of a male Dutch Flemish voice. Extracted from Mozilla Common Voice project [https://github.com/common-voice/common-voice/tree/master/server/data/nl].
mozilla_commonvoice_Arabic_preprocessed_train_batch_1pexels-gpt4oImages collected from Pexels, using 1000 images following 5 categories:
nudes
war
group
animals
smoking
See https://www.pexels.com/license/ for the license
They were then annotated using gpt4-o, see https://github.com/mozilla/distilvit/blob/main/distilvit/gpt4.py
mozilla-common-voice-converted-to-parquet-pt
