datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
demo1
Dataset Card for Demo1
Dataset Summary
This is a demo dataset. It consists in two files data/train.csv and data/test.csv
You can load it with
from datasets import load_dataset
demo1 = load_dataset("lhoestq/demo1")
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/lhoestq/demo1.fava-flagged-demo
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/abhika-m/fava-flagged-demo.crowdsourced-calculator-demohdi-ihdi-democracy-by-countryData sourced from:
United Nations Development Program:
https://hdr.undp.org/sites/default/files/2023-24_HDR/HDR23-24_Statistical_Annex_I-HDI_Table.xlsx
Economist Intelligence Unit Democracy Index through Our World In Data:
https://ourworldindata.org/grapher/democracy-index-eiu.csv?v=1&csvType=full&useColumnShortNames=true
There is missing data of course, since both of the reports for 2023 is limited, so you can decide on your own how you want to filter.
I might later post a backfilled variant… See the full description on the dataset page: https://huggingface.co/datasets/marksverdhei/hdi-ihdi-democracy-by-country.demodiff_downstream
DemoDiff Downstream Context Data
This dataset contains context data for the DemoDiff project, a diffusion-based molecular foundation model for in-context inverse molecular design. It provides contextual examples to guide molecular generation, enabling few-shot molecular design across diverse chemical tasks without task-specific fine-tuning.
Structure
Each task is organized as a separate folder in the repository root:
tasks/
├── Albuterol_Similarity/
│ ├── positive.csv
│… See the full description on the dataset page: https://huggingface.co/datasets/liuganghuggingface/demodiff_downstream.ArchEGraph-demo
ArchEGraph-demo
ArchEGraph-demo is a compact demo package of the ArchEGraph building-energy dataset for graph-based and weather-conditioned learning.
Dataset Summary
Total cases in manifest.csv: 300
Unique buildings: 75
Unique weather IDs: 48
n_steps: always 8,760
n_spaces range: 2 to 132
This package currently stores:
manifest.csv (index of all demo cases)
building/ (75 files)
geometry/ (75 files)
weather/ (48 files)
energy/ (300 files)
split/ (demo split CSV files)… See the full description on the dataset page: https://huggingface.co/datasets/ArchEGraph/ArchEGraph-demo.mozart-api-demo-pages
Dataset Card for Dataset Name
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/DoctorSlimm/mozart-api-demo-pages.demo2
Dataset Card for Demo1
Dataset Summary
This is a demo dataset. It consists in two files data/train.csv and data/test.csv
You can load it with
from datasets import load_dataset
demo1 = load_dataset("lhoestq/demo1")
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/lhoestq/demo2.gc_santegc_sante (Grande Cause Santé; Health Great Cause) is a French citizen consultation on how to act collectively for better health, prevention an well-being in France held in 2025-2026.
Data
This dataset contains three subsets:
proposals contains the written proposals in French. Each proposal has a text content and has a unique id proposal_id. the topic and subtopic column contains the LLM-generated and human-validated topics (clusters) used during the official analysis of the… See the full description on the dataset page: https://huggingface.co/datasets/democratic-commons/gc_sante.image-preference-demo
Image dataset for preference aquisition demo
This dataset provides the files used to run the example that we use in this blog post to illustrate how easily
you can set up and run the annotation process to collect a huge preference dataset using Rapidata's API.
The goal is to collect human preferences based on pairwise image matchups.
The dataset contains:
Generated images: A selection of example images generated using Flux.1 and Stable Diffusion. The images are provided in a .zip… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/image-preference-demo.ingerenceIngerence is a French citizen consultation on how to combat information manipulation due to foreign digital interference held in 2025-2026.
Data
This dataset contains three subsets:
proposals contains the written proposals in French. Each proposal has a text content written by an author with a unique author_id, and has a unique id proposal_id.
votes contains the votes of users on propositions. Each user has a unique id user_id and votes on proposals (defined by proposal_id).… See the full description on the dataset page: https://huggingface.co/datasets/democratic-commons/ingerence.crowdsourced-calculator-demoreddit-comments-demoeurhopeEurhope is a European-Union wide multilingual citizens consultation on the future of the European Union. It was held in 2024.
Data
This dataset contains two subsets:
proposals contains the written proposals in French. Each proposal has a text content written by a author author_id and has a unique id proposal_id. Each proposal has one of 22 language given in its language column.
votes contains the votes of users on propositions. Each user has a unique id user_id and votes on… See the full description on the dataset page: https://huggingface.co/datasets/democratic-commons/eurhope.DemosQA
DemosQA
We introduce DemosQA (δῆμος), a novel Greek QA dataset, which is constructed using social media user questions and community-reviewed answers to better capture the Greek social and cultural zeitgeist.
It comprises questions extracted from the “r/greece” subreddit, each accompanied by four candidate answers, the selected best answer and its index, the date of posting, and the corresponding Reddit post ID.
Candidate answers are ranked based on community voting, with the… See the full description on the dataset page: https://huggingface.co/datasets/IMISLab/DemosQA.steuer_debateThe Steuer Debate consultation is a german citizen participation project on fair taxes and finances held in 2025.
Data
This dataset contains three subsets:
proposals contains the written propositions (in german). the topic column contains the LLM-generated and human-validated topics (clusters) used during the official analysis of the consultation. Each proposal has a unique id proposal_id.
votes contains the votes of users on propositions. Each user has a unique id user_id and… See the full description on the dataset page: https://huggingface.co/datasets/democratic-commons/steuer_debate.math-intuition-20260906-403-demo-10
math-intuition-20260906-403-demo-10
3,936 mathematics problems drawn from 403 problem families, each derived from a
distinct arXiv paper. Every problem is generated answer-first, so the answer is known by
construction and is checked by the family's own verify() before the row is written.
No row in this file is ungraded.
This is the demo rung — read this before using it
Each family exposes a four-rung ladder: demo, easy, medium, hard. This file samples
demo, which… See the full description on the dataset page: https://huggingface.co/datasets/amphora/math-intuition-20260906-403-demo-10.eval-harness-demo
OmicsBank Eval Harness (Demo)
⚠️ This uses entirely synthetic data. No row corresponds to a real patient. Built to
demonstrate the harness mechanics — not to represent real-world statistics or actual
clinical patterns.
Two tools, both runnable in a couple minutes with just pandas installed:
Auto-grader — score your model's predictions on a small held-out classification task
Leakage checker — check whether your eval/training data overlaps with a reference corpus… See the full description on the dataset page: https://huggingface.co/datasets/omicsbank/eval-harness-demo.demo-salaries
Dataset Summary
Briefly summarize the dataset, its intended use and the supported tasks. Give an overview of how and why the dataset was created. The summary should explicitly mention the languages present in the dataset (possibly in broad terms, e.g. translations between several pairs of European languages), and describe the domain, topic, or genre covered.
Supported Tasks and Leaderboards
For each of the tasks tagged for this dataset, give a brief description of the tag… See the full description on the dataset page: https://huggingface.co/datasets/Einstellung/demo-salaries.Law-Demographic-Bias-Difference-Awareness
Law and Demographic Bias Difference-Awareness Benchmark
A multiple-choice benchmark for testing whether a language model can tell apart two situations
that look alike and demand opposite answers:
neq — the law grants an entitlement to one specific group, so treating both groups
identically is the wrong answer.
eq — the law grants the same right to everyone, so drawing a distinction between the
groups is the wrong answer.
Every item presents two demographic or legal groups, a… See the full description on the dataset page: https://huggingface.co/datasets/Debk/Law-Demographic-Bias-Difference-Awareness.student-grades-demo
Student Grades Demo Dataset
This dataset contains student grades data with both true labels and noisy (corrupted) labels.
Dataset Description
The dataset includes:
Student exam scores (exam_1, exam_2, exam_3)
Notes field
True letter grades (letter_grade)
Noisy/corrupted letter grades (noisy_letter_grade)
This is useful for demonstrating and validating label error detection methods.
Usage
import pandas as pd
# Load the dataset
df =… See the full description on the dataset page: https://huggingface.co/datasets/Cleanlab/student-grades-demo.wmdp-defense-demoek100-mir-demo-assetsDemo_Solubilitydiabetes
Dataset Card for Auditor_Review
This file is a copy, the original version is hosted at data.world
demodbenchmark-demo-results
Benchmark Demo Results
This public dataset stores demo leaderboard results for the Hugging Face leaderboard MVP.
The main table is leaderboard.csv. In the MVP, the metrics are demo values generated from an example submission. Later this file will be updated by the Gradio Space after real submissions are scored.
crowdsourced-calculator-demo
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/adonaivera/crowdsourced-calculator-demo.africa-synth-demographics-demographic-projections-all
African Demographic Projections (2020-2050) | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: csv - Sector: demographics_social - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-demographics-demographic-projections-all.flipkart-preprocessed-demo
https://github.com/GoogleCloudPlatform/accelerated-platforms/tree/main/docs/use-cases/model-fine-tuning-pipeline#data-preprocessing-steps
https://www.kaggle.com/datasets/PromptCloudHQ/flipkart-products/data
