datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
commonsense_qa
Dataset Card for "commonsense_qa"
Dataset Summary
CommonsenseQA is a new multiple-choice question answering dataset that requires different types of commonsense knowledge
to predict the correct answers . It contains 12,102 questions with one correct answer and four distractor answers.
The dataset is provided in two major training/validation/testing set splits: "Random split" which is the main evaluation
split, and "Question token split", see paper for details.… See the full description on the dataset page: https://huggingface.co/datasets/tau/commonsense_qa.commonsense-qaYouTube-Commons
📺 YouTube-Commons 📺
YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC-By license.
Content
The collection comprises 22,709,724 original and automatically translated transcripts from 3,156,703 videos (721,136 individual channels).
In total, this represents nearly 45 billion words (44,811,518,375).
All the videos where shared on YouTube with a CC-BY license: the dataset provide all the necessary provenance information… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/YouTube-Commons.compact-alignments
compact-alignments — per-verse, per-book, content-addressed
The token-position companion to lexeme-alignments (which is
aggregated/type-level and can't tell you what happened in any one verse). This dataset restores
position: for a given edition's Bible book, which Hebrew/Greek content word aligned to which
target-text token, verse by verse.
The authoritative list of what's published is always manifest.json, not this file.
Original-language source editions (needed… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/compact-alignments.german-commons
German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
A comprehensive collection of German-language text data under open licenses for training German language models.
Datasheet: DATASHEET.md.
Paper: arxiv.org/abs/2510.13996
Code: github.com/coral-nlp/llmdata
Bloom Filter (DOLMA-compatible): bloom_filter.bin
Dataset Description
This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/coral-nlp/german-commons.german-commons
German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
A comprehensive collection of German-language text data under open licenses for training German language models.
Datasheet: DATASHEET.md.
Paper: arxiv.org/abs/2510.13996
Code: github.com/coral-nlp/llmdata
Bloom Filter (DOLMA-compatible): bloom_filter.bin
Dataset Description
This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/CUI03/german-commons.Medical-Commons
Medical-Commons
Medical-Commons is the largest dataset of medical content under free licenses or open data program collected by Pleias.
It includes three different collection:
International scientific collection of 2M articles from OpenAlex.
French scientific collection of XM articles, reports and PhD theses from French institutional repositories.
Administration collection from health and medical agencies, for now limited to France but with a planned Europe-wide expansion.
The… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Medical-Commons.agent-traces
Trace Commons — Agent Traces
Trace Commons is one open, public dataset of coding-agent sessions — the
back-and-forth between a developer and an AI coding agent, including prompts,
model responses, tool calls, and command output — contributed voluntarily as an
open resource for studying, evaluating, and building on how these agents
actually work.
Every trace here was donated only from a public, open-source repository, was
anonymized on the contributor's own machine before upload… See the full description on the dataset page: https://huggingface.co/datasets/trace-commons/agent-traces.OpenEval
OpenEval
An open-source, item-centered evaluation repository toward the open science of AI evaluation.
This official dataset is maintained by the Open Eval Commons project.
For using or contributing to OpenEval (thank you!), please refer to our detailed documentation.
🌐 OpenEval Homepage | 📦 GitHub Repository
🏗️Dataset Structure
Currently, the data are split into three tables for storage efficiency:
bench, where bench entries are indexed by the field… See the full description on the dataset page: https://huggingface.co/datasets/Open-Eval-Commons/OpenEval.physical-commonsensecommonsense_qa_2.0https://github.com/allenai/csqa2
@article{talmor2022commonsenseqa,
title={CommonsenseQA 2.0: Exposing the limits of AI through gamification},
author={Talmor, Alon and Yoran, Ori and Bras, Ronan Le and Bhagavatula, Chandra and Goldberg, Yoav and Choi, Yejin and Berant, Jonathan},
journal={arXiv preprint arXiv:2201.05320},
year={2022}
}
CommonSketch
CommonSketch
Dataset Summary
CommonSketch is a semantically annotated sketch dataset introduced in the paper SEA: Evaluating Sketch Abstraction Efficiency via Element-level Commonsense Visual Question Answering. The dataset contains 23,100 human-drawn sketches across 300 object classes. Each sketch is paired with a fine-grained caption and element-level commonsense annotations for evaluating sketch abstraction and semantic recognizability.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/ziiio/CommonSketch.French-Science-Commons
French Science Commons
French Science Commons (Commun numérique des sciences en français) rassemble des publications scientifiques d'origine française en accès ouvert, couvrant une période de vingt ans, de 2007 à 2026. Il comprend 1 248 860 documents scientifiques — 1 189 628 articles et 59 232 thèses — indexés à travers de multiples dépôts académiques en accès public, tels que HAL, OpenAlex, des revues scientifiques, des dépôts institutionnels, et d'autres.
Le corpus est conçu… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-Science-Commons.DemoFeedbackaligned-mwe
aligned_mwe — multi-word target expressions per lexeme
Where lexeme-alignments is one row per surface
token, this is one row per lexeme rendered by a contiguous multi-word phrase (חֶסֶד → "kasih setia",
בֵּית → "tempat pengirikan"). Mined from the aligner's per-verse t_idx positions: only spans whose
target token positions are contiguous (max−min+1 == len) qualify — scattered tokens that merely all
linked to a lexeme are dropped (and counted in the manifest as scattered_dropped… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/aligned-mwe.target-stopwords
target-stopwords
Per-language function-word lists, induced from that language's own Bible text — frequency +
dispersion (the classic corpus-linguistics stopword-induction recipe), then RESCUED against the
language's own alignment output + a source-anchored content signal so genuinely frequent CONTENT words
("God", "Lord") are never dropped.
A candidate word is rescued out of the list (judged a real content word, not a function word) only when
all four hold — see the… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/target-stopwords.wikimedia-commons-documents-ml_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_beir.commonsense_170khttps://github.com/AGI-Edgerunners/LLM-Adapters/blob/main/ft-training_set/commonsense_170k.json
target-morphology
target-morphology
Per-language unsupervised morphology models — productive suffixes, prefixes, and a stem lexicon,
each learned MDL-free ("Linguistica"-style: a suffix is productive if it attaches to many paradigm stems)
from that language's own Bible text. No labels, no pretrained model, no download — so it runs on any
language with a translation, including those with zero LLM/encoder coverage.
stem(word) strips one productive affix when the remainder is a known stem; inflected… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/target-morphology.senses-attested
senses_attested — attested target renderings per lexeme sense
The empirical evidence layer produced for shoresh (bcv-query data-contract): for a lexeme in a
disambiguated (binyan, sense), which target-language words attest it, with counts. It is the supply
that fills shoresh's senses_i18n/_gaps demand and cross-checks the llm_strongs_glosses predictions —
it does not replace shoresh's curated senses_i18n/<iso>.tsv; consumed as an HF Parquet dataset.
Schema (per row)… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/senses-attested.YouTube-Commons
📺 YouTube-Commons 📺
YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC-By license.
Content
The collection comprises 22,709,724 original and automatically translated transcripts from 3,156,703 videos (721,136 individual channels).
In total, this represents nearly 45 billion words (44,811,518,375).
All the videos where shared on YouTube with a CC-BY license: the dataset provide all the necessary provenance… See the full description on the dataset page: https://huggingface.co/datasets/gautijha37/YouTube-Commons.EU-Science-Commonslexeme-alignments
lexeme-alignments — surface → original-language lexeme (Strong's-bridged)
For each language, the attested mapping from target surface word-forms → the original-language
lexeme they render, mined by the aligner. Lexeme-anchored, provenance-honest, additive — the
design principles are in docs/publishing-principles.md. One
language per partition, for consumption by bcv-commons and downstream tools.
The language: list above tracks the published partitions; the authoritative list is… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/lexeme-alignments.task828_copa_commonsense_cause_effect
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task828_copa_commonsense_cause_effect
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task828_copa_commonsense_cause_effect.wikimedia-commons-maps_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-maps reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-maps_beir.YouTube-Commons-nl-audio
YouTube Commons NL Audio
This dataset has audio files for the Dutch-language videos in Rijgersberg/YouTube-Commons-nl-transcriptions,
all under a CC BY 4.0 license.
It contains 11,669 files for a total runtime of 2493h 43m 5s, coming in at approximately 130 GB.
Source
The original source of the dataset (minus the titles, descriptions and audio files) is YouTube Commons:
YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons-nl-audio.common_sense_reasoningThis is the repository containing LoRA checkpoints tuned on common sense reasoning datasets with 0.5B foundation model, served as training data for DnD.
safe-commons-pd-3m
Safe Commons PD 3M
This is a balanced and safe-to-use public domain / CC0 images dataset.
All images and texts come from Wikimedia Commons and Wikidata with strict filtering.
Images license is either Public Domain or CC0 (varies by image).
Texts license is either CC0 or CC BY-SA (varies by caption source).
No synthetic data (AI generated images or captions) is in the dataset.
To build this dataset, we tried to avoid any knowledge leaks from existing pre-trained models at the… See the full description on the dataset page: https://huggingface.co/datasets/Mitsua/safe-commons-pd-3m.vel_commons_wikidata
Visual Entity Linking: Wikimedia Commons & Wikidata
This dataset allows to train and evaluate ML models that link Wikimedia Commons images to the Wikidata items they depict.
Disclaimer: All images contained in this dataset are generally assumed to be freely usable (as intended for Wikimedia Commons). Each image's license and author/
uploader is - to the best of our ability - reported in its metadata (see section Dataset Structure). If you want your image's attribution changed or the… See the full description on the dataset page: https://huggingface.co/datasets/aiintelligentsystems/vel_commons_wikidata.YouTube-Commons
YouTube Commons Re-upload
This is a re-upload of PleIAs' YouTube Commons, a valuable open dataset:
YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC BY 4.0 license.
Content
The collection comprises 22,709,724 original and automatically translated transcripts from 3,156,703 videos (721,136 individual channels).
Unfortunately, there are problems with loading YouTube Commons with Hugging Face Datasets.
In order to alleviate those… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons.
