datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
demo1
Dataset Card for Demo1
Dataset Summary
This is a demo dataset. It consists in two files data/train.csv and data/test.csv
You can load it with
from datasets import load_dataset
demo1 = load_dataset("lhoestq/demo1")
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/lhoestq/demo1.fava-flagged-demo
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/abhika-m/fava-flagged-demo.crowdsourced-calculator-demohdi-ihdi-democracy-by-countryData sourced from:
United Nations Development Program:
https://hdr.undp.org/sites/default/files/2023-24_HDR/HDR23-24_Statistical_Annex_I-HDI_Table.xlsx
Economist Intelligence Unit Democracy Index through Our World In Data:
https://ourworldindata.org/grapher/democracy-index-eiu.csv?v=1&csvType=full&useColumnShortNames=true
There is missing data of course, since both of the reports for 2023 is limited, so you can decide on your own how you want to filter.
I might later post a backfilled variant… See the full description on the dataset page: https://huggingface.co/datasets/marksverdhei/hdi-ihdi-democracy-by-country.skill-demand-index
Datamata Skill Demand Index
Daily share of active tech job listings mentioning each skill, across data, engineering, product, DevOps, security and AI. One row per category and skill from the most recent snapshot, including how often each skill is a hard requirement.
Latest snapshot: 2026-09-22
Rows in this release: 753
Updated: daily
Licence: CC BY 4.0 — free to use and adapt, including commercially, with attribution.
Source & methodology:… See the full description on the dataset page: https://huggingface.co/datasets/datamatastudios/skill-demand-index.demodiff_downstream
DemoDiff Downstream Context Data
This dataset contains context data for the DemoDiff project, a diffusion-based molecular foundation model for in-context inverse molecular design. It provides contextual examples to guide molecular generation, enabling few-shot molecular design across diverse chemical tasks without task-specific fine-tuning.
Structure
Each task is organized as a separate folder in the repository root:
tasks/
├── Albuterol_Similarity/
│ ├── positive.csv
│… See the full description on the dataset page: https://huggingface.co/datasets/liuganghuggingface/demodiff_downstream.traffic-demand-csv-files-1dementor-sft-dataArchEGraph-demo
ArchEGraph-demo
ArchEGraph-demo is a compact demo package of the ArchEGraph building-energy dataset for graph-based and weather-conditioned learning.
Dataset Summary
Total cases in manifest.csv: 300
Unique buildings: 75
Unique weather IDs: 48
n_steps: always 8,760
n_spaces range: 2 to 132
This package currently stores:
manifest.csv (index of all demo cases)
building/ (75 files)
geometry/ (75 files)
weather/ (48 files)
energy/ (300 files)
split/ (demo split CSV files)… See the full description on the dataset page: https://huggingface.co/datasets/ArchEGraph/ArchEGraph-demo.Demacia
DeepSearchQA
A 900-prompt factuality benchmark from Google DeepMind, designed to evaluate agents on difficult multi-step information-seeking tasks across 17 different fields.
▶ Google DeepMind Release Blog Post▶ DeepSearchQA Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code
Benchmark
DeepSearchQA is a 900-prompt benchmark for evaluating agents on difficult multi-step information-seeking tasks across 17 different fields. Unlike traditional… See the full description on the dataset page: https://huggingface.co/datasets/Ryanduy/Demacia.mozart-api-demo-pages
Dataset Card for Dataset Name
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/DoctorSlimm/mozart-api-demo-pages.US_Presidential_Election_2020_Dem_Repdemo2
Dataset Card for Demo1
Dataset Summary
This is a demo dataset. It consists in two files data/train.csv and data/test.csv
You can load it with
from datasets import load_dataset
demo1 = load_dataset("lhoestq/demo1")
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/lhoestq/demo2.dementor-matrix-responses
Dementor — matrix model responses
Generated model outputs for the Dementor LLM-imitation / behavioral-inertia study.
Companion to:
Code + prompt splits: https://github.com/lisadunlap/dementor (branch ethan)
Trained adapters (2,122 LoRAs): https://huggingface.co/dementor-research — SFT / DPO /
self-SFT, grouped into per-dataset collections (gsm8k, chatbot_arena, writingprompts, openassistant).
Dataset viewer. This repo is a nested tree of CSV tables plus per-cell cell.json… See the full description on the dataset page: https://huggingface.co/datasets/dementor-research/dementor-matrix-responses.gc_santegc_sante (Grande Cause Santé; Health Great Cause) is a French citizen consultation on how to act collectively for better health, prevention an well-being in France held in 2025-2026.
Data
This dataset contains three subsets:
proposals contains the written proposals in French. Each proposal has a text content and has a unique id proposal_id. the topic and subtopic column contains the LLM-generated and human-validated topics (clusters) used during the official analysis of the… See the full description on the dataset page: https://huggingface.co/datasets/democratic-commons/gc_sante.image-preference-demo
Image dataset for preference aquisition demo
This dataset provides the files used to run the example that we use in this blog post to illustrate how easily
you can set up and run the annotation process to collect a huge preference dataset using Rapidata's API.
The goal is to collect human preferences based on pairwise image matchups.
The dataset contains:
Generated images: A selection of example images generated using Flux.1 and Stable Diffusion. The images are provided in a .zip… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/image-preference-demo.ingerenceIngerence is a French citizen consultation on how to combat information manipulation due to foreign digital interference held in 2025-2026.
Data
This dataset contains three subsets:
proposals contains the written proposals in French. Each proposal has a text content written by an author with a unique author_id, and has a unique id proposal_id.
votes contains the votes of users on propositions. Each user has a unique id user_id and votes on proposals (defined by proposal_id).… See the full description on the dataset page: https://huggingface.co/datasets/democratic-commons/ingerence.dementor-matrix-baselinescrowdsourced-calculator-demoreddit-comments-demoeurhopeEurhope is a European-Union wide multilingual citizens consultation on the future of the European Union. It was held in 2024.
Data
This dataset contains two subsets:
proposals contains the written proposals in French. Each proposal has a text content written by a author author_id and has a unique id proposal_id. Each proposal has one of 22 language given in its language column.
votes contains the votes of users on propositions. Each user has a unique id user_id and votes on… See the full description on the dataset page: https://huggingface.co/datasets/democratic-commons/eurhope.DemosQA
DemosQA
We introduce DemosQA (δῆμος), a novel Greek QA dataset, which is constructed using social media user questions and community-reviewed answers to better capture the Greek social and cultural zeitgeist.
It comprises questions extracted from the “r/greece” subreddit, each accompanied by four candidate answers, the selected best answer and its index, the date of posting, and the corresponding Reddit post ID.
Candidate answers are ranked based on community voting, with the… See the full description on the dataset page: https://huggingface.co/datasets/IMISLab/DemosQA.dementor-steering-directionssteuer_debateThe Steuer Debate consultation is a german citizen participation project on fair taxes and finances held in 2025.
Data
This dataset contains three subsets:
proposals contains the written propositions (in german). the topic column contains the LLM-generated and human-validated topics (clusters) used during the official analysis of the consultation. Each proposal has a unique id proposal_id.
votes contains the votes of users on propositions. Each user has a unique id user_id and… See the full description on the dataset page: https://huggingface.co/datasets/democratic-commons/steuer_debate.math-intuition-20260906-403-demo-10
math-intuition-20260906-403-demo-10
3,936 mathematics problems drawn from 403 problem families, each derived from a
distinct arXiv paper. Every problem is generated answer-first, so the answer is known by
construction and is checked by the family's own verify() before the row is written.
No row in this file is ungraded.
This is the demo rung — read this before using it
Each family exposes a four-rung ladder: demo, easy, medium, hard. This file samples
demo, which… See the full description on the dataset page: https://huggingface.co/datasets/amphora/math-intuition-20260906-403-demo-10.clinical-quad-oxygen-demand-buffer-lag-coupling-respiratory-collapse-v0.6
What this repo does
This repository contains a Clarus v0.6 intervention pathway dataset focused on respiratory collapse dynamics.
The dataset evaluates whether a model can determine if a proposed intervention meaningfully stabilizes a deteriorating respiratory system.
The task requires reasoning from:
system state
trajectory toward instability
boundary geometry
recovery geometry
intervention vector
projected trajectory consequence
The model cannot read the answer directly.
It must… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-oxygen-demand-buffer-lag-coupling-respiratory-collapse-v0.6.eval-harness-demo
OmicsBank Eval Harness (Demo)
⚠️ This uses entirely synthetic data. No row corresponds to a real patient. Built to
demonstrate the harness mechanics — not to represent real-world statistics or actual
clinical patterns.
Two tools, both runnable in a couple minutes with just pandas installed:
Auto-grader — score your model's predictions on a small held-out classification task
Leakage checker — check whether your eval/training data overlaps with a reference corpus… See the full description on the dataset page: https://huggingface.co/datasets/omicsbank/eval-harness-demo.demo-salaries
Dataset Summary
Briefly summarize the dataset, its intended use and the supported tasks. Give an overview of how and why the dataset was created. The summary should explicitly mention the languages present in the dataset (possibly in broad terms, e.g. translations between several pairs of European languages), and describe the domain, topic, or genre covered.
Supported Tasks and Leaderboards
For each of the tasks tagged for this dataset, give a brief description of the tag… See the full description on the dataset page: https://huggingface.co/datasets/Einstellung/demo-salaries.most-in-demand-skills-2026
Most In-Demand Job Skills of 2026
Skill-demand frequencies extracted from 360,000+ job postings collected by Qarera between December 27, 2025 and June 16, 2026.
📊 Full report & charts: The Most In-Demand Skills of 2026
🔖 Cite this dataset (DOI): 10.5281/zenodo.21204423
📄 License: CC BY 4.0 — free to use with attribution to Qarera.
Key findings
We counted the skills named in 360,000+ job postings (Dec 2025–Jun 2026).
"AI" was the #2 most-requested skill overall… See the full description on the dataset page: https://huggingface.co/datasets/yash2111/most-in-demand-skills-2026.demandpulse
