datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CommonsenseQA-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in CommonsenseQA. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
gc_santegc_sante (Grande Cause Santé; Health Great Cause) is a French citizen consultation on how to act collectively for better health, prevention an well-being in France held in 2025-2026.
Data
This dataset contains three subsets:
proposals contains the written proposals in French. Each proposal has a text content and has a unique id proposal_id. the topic and subtopic column contains the LLM-generated and human-validated topics (clusters) used during the official analysis of the… See the full description on the dataset page: https://huggingface.co/datasets/democratic-commons/gc_sante.ingerenceIngerence is a French citizen consultation on how to combat information manipulation due to foreign digital interference held in 2025-2026.
Data
This dataset contains three subsets:
proposals contains the written proposals in French. Each proposal has a text content written by an author with a unique author_id, and has a unique id proposal_id.
votes contains the votes of users on propositions. Each user has a unique id user_id and votes on proposals (defined by proposal_id).… See the full description on the dataset page: https://huggingface.co/datasets/democratic-commons/ingerence.eurhopeEurhope is a European-Union wide multilingual citizens consultation on the future of the European Union. It was held in 2024.
Data
This dataset contains two subsets:
proposals contains the written proposals in French. Each proposal has a text content written by a author author_id and has a unique id proposal_id. Each proposal has one of 22 language given in its language column.
votes contains the votes of users on propositions. Each user has a unique id user_id and votes on… See the full description on the dataset page: https://huggingface.co/datasets/democratic-commons/eurhope.steuer_debateThe Steuer Debate consultation is a german citizen participation project on fair taxes and finances held in 2025.
Data
This dataset contains three subsets:
proposals contains the written propositions (in german). the topic column contains the LLM-generated and human-validated topics (clusters) used during the official analysis of the consultation. Each proposal has a unique id proposal_id.
votes contains the votes of users on propositions. Each user has a unique id user_id and… See the full description on the dataset page: https://huggingface.co/datasets/democratic-commons/steuer_debate.hebrew-lexical-references
Hebrew Lexical Reference Indices
Four structured, Strong's-linked transcriptions of external Hebrew (and one Hebrew↔Greek) lexical
reference sources. These are not our own synonymy judgments — each config faithfully represents
what an established outside source, or an actual historical translation record, already asserts (an
etymological dictionary's own root groupings, a WordNet's own synset membership, five named scholars'
own verified structural analysis, the Septuagint's own… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/hebrew-lexical-references.common-strategy-d0489c
common-strategy-d0489c
Synthetic sensors test data: 46 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Ember-Wisp/common-strategy-d0489c.philosophy-culture-translations-html-csv
AI-Culture Philosophy and Culture Translations CSV + HTML Corpus
The corpus contains an exceptionally diverse range of cultural, philosophical, and literary texts, available in 12 major languages. Among other topics, there is extensive engagement with the ethics and aesthetics of artificial intelligence and its cultural and philosophical implications, as well as connections between AI and philosophy of language and philosophy of mind.
This project is maintained by a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/philosophy-culture-translations-html-csv.commonSenseWithExplicationscommonSenseWithExplicationsAndEmbeddingcommonsense_dialogues
