datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sweden_100K_difficultccs_dataset_summarised_diffdiffssdBelow, we present the license requirements incorporated by reference and README, which explains the dataset and directory structure. Both sections have different
formatting to help with easy navigation.
=======================================================================
DiffSSD: Diffusion-based Synthetic Speech Dataset
=======================================================================
The DiffSSD (Diffusion-based Synthetic Speech Dataset) has been derived
Using real speech signals… See the full description on the dataset page: https://huggingface.co/datasets/purdueviperlab/diffssd.diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04 Contains maximum activating examples for all the features of our crosscoder trained on gemma 2 2B layer 13 available here: https://huggingface.co/Butanium/gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04/blob/main/README.md
base_examples.pt contains all the maximum examples of the feature on a subset of validation test of fineweb
chat_examples.pt is the same but for lmsys chat data
chat_base_examples.pt is a merge of the two above files.
All files are of the type dict[int, list[tuple[float… See the full description on the dataset page: https://huggingface.co/datasets/science-of-finetuning/diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04.diffusion-mcqa-gen-pelatnas-2026
Which Prompt Made This? — Generated Edition
Pelatnas IOAI 2026 · Task Diffusion MCQA (varian trajectory)
Sebuah model text-to-image sedang bekerja. Di tengah prosesnya, gambar belum
menjadi gambar — yang ada hanya latent ter-noise: tensor 4 × 64 × 64 berisi
campuran struktur yang mulai muncul dan derau Gaussian.
Kali ini kalimat itu harfiah. Latent yang kamu terima benar-benar diambil dari
tengah proses generate: sebuah trajectory denoising DDIM 50 langkah dihentikan
sejenak… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/diffusion-mcqa-gen-pelatnas-2026.benchmarks
Welcome to 🤗 Diffusers Benchmarks!
This is dataset where we keep track of the inference latency and memory information of the core models in the diffusers library.
Currently, the core models are:
Flux
Wan
LTX
SDXL
Note that we will continue to extend this list based on their usage.
You can analyze the results in this demo.
[!IMPORTANT]
Instead of benchmarking the entire diffusion pipelines, we only benchmark the forward passes
of the diffusion networks under different settings… See the full description on the dataset page: https://huggingface.co/datasets/diffusers/benchmarks.git-diff_to_commit_msg
Hi, I’m Seniru Epasinghe 👋
I’m an AI undergraduate and an AI enthusiast, working on machine learning projects and open-source contributions.I enjoy exploring AI pipelines, natural language processing, and building tools that make development easier.
🌐 Connect with me
There are 2 version of this dataset:
git-diff_to_commit_msg - 1.5K rows
huggingface link
kaggle link
git-diff_to_commit_msg_large - 1.75M rows
huggingface link
kaggle link… See the full description on the dataset page: https://huggingface.co/datasets/seniruk/git-diff_to_commit_msg.latent-dna-diffusionDeepScaleR_Difficulty
Difficulty Estimation on DeepScaleR
We annotate the entire DeepScaleR dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation.
DeepScaleR is a curated dataset of 40,000 reasoning-intensive problems used to train and evaluate reinforcement learning-based methods for large language models.
Difficulty Scoring Method
Difficulty scores are estimated using the… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/DeepScaleR_Difficulty.Diffusion-Reward-Modeling-for-Text-Rendering-Dataset
🖼️ Text-to-Image Rendering Dataset
A dataset of 14k text prompts for image generation with text rendering evaluation
📚 Dataset Overview
This dataset contains 14,000 text prompts specifically designed for:
Image generation with text rendering
Evaluating text preservation in generated images
Training diffusion models for better text rendering
Each prompt comes with:
Pre-extracted target text for rendering
5 Stable Diffusion 3 generated latents (70k total)
Dual… See the full description on the dataset page: https://huggingface.co/datasets/leffff/Diffusion-Reward-Modeling-for-Text-Rendering-Dataset.MMA-Diffusion-NSFW-adv-prompts-benchmark
MMA-Diffusion Adversarial Prompts (Text modal attack)
The MMA-Diffusion adversarial prompts benchmark comprises 1,000 successful adversarial prompts generated by the adversarial attack methodology presented in the paper
from CVPR 2024 titled MMA-Diffusion: MultiModal Attack on Diffusion Models. This resource is intended to assist in developing and
evaluating defense mechanisms against such attacks. The adversarial prompts are capable of bypassing the image safety checker in… See the full description on the dataset page: https://huggingface.co/datasets/YijunYang280/MMA-Diffusion-NSFW-adv-prompts-benchmark.diffusion-vs-ar-hard-sudoku
Diffusion vs AR Hard Sudoku
This repository packages 8,148,696 Sudoku examples in the CSV format expected
by HKUNLP/diffusion-vs-ar, plus its original 100k/1k easy baseline.
Every processed file has these columns:
column
meaning
quizzes
81 row-major digits; 0 is an empty cell
solutions
complete 81-digit solution
source
original collection
dataset
normalized dataset family
official_rating
rating supplied by the source
rating_type
semantics of that rating… See the full description on the dataset page: https://huggingface.co/datasets/fhyfhy/diffusion-vs-ar-hard-sudoku.GSM8K_Difficulty
Difficulty Estimation on DeepScaleR
We annotate the entire GSM8K dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation.
GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/GSM8K_Difficulty.human_anatomy_qa_with_difficulty
Truth, Trust, and Trouble (TTT) – Medical Anatomy QA Benchmark
This repository hosts the dataset introduced in the EMNLP Industry Track 2025 paper “Truth, Trust, and Trouble: Medical AI on the Edge.”
The dataset contains 1,077 high-quality, clinically validated True/False anatomy questions, designed to evaluate medical LLMs along three critical axes:
Honesty (factual alignment)
Helpfulness (semantic relevance & completeness)
Harmlessness (safety under clinical constraints)
This… See the full description on the dataset page: https://huggingface.co/datasets/ekplatebiryani/human_anatomy_qa_with_difficulty.git-diff_to_commit_msg_large
Hi, I’m Seniru Epasinghe 👋
I’m an AI undergraduate and an AI enthusiast, working on machine learning projects and open-source contributions.I enjoy exploring AI pipelines, natural language processing, and building tools that make development easier.
🌐 Connect with me
There are 2 version of this dataset:
git-diff_to_commit_msg - 1.5K rows
huggingface link
kaggle link
git-diff_to_commit_msg_large - 1.75M rows
huggingface link
kaggle link… See the full description on the dataset page: https://huggingface.co/datasets/seniruk/git-diff_to_commit_msg_large.orz_math_difficulty
Difficulty Estimation on Open Reasoner Zero
We annotate the entire Open Reasoner Zero dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction.
Open Reasoner Zero is a curated a dataset of 57,000 reasoning-intensive problems used to train and evaluate reinforcement learning-based methods for large language models.
Difficulty Scoring Method
Difficulty scores are estimated using… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/orz_math_difficulty.differential-preservation-narrative-v01Differential Preservation Under Narrative PressureClinical Narrative Integrity v0.2
Purpose
Test whether models preserve multiple plausible diagnoses
Test whether narrative fluency collapses uncertainty
Test resistance to premature diagnostic closure
Central question
What else could this be
Why this dataset exists
Narrative pressure rewards coherence.Clinical safety requires openness.
This dataset isolates the moment where a single story becomes dominant despite nonspecific evidence.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/differential-preservation-narrative-v01.stable_diffusion_prompts_instruct
Stable diffusion prompts for instruction models fine-tuning
Overview
This dataset contains 80,000+ prompts summarized to make it easier to create instruction-tuned prompt enhancing models. Each row of the dataset contains two values:
a short description of a image
a full prompt corresponding to that description in a stable diffusion format
Hope this dataset can help creating amazing apps !
How to use
You can download and use the dataset easily using the… See the full description on the dataset page: https://huggingface.co/datasets/groloch/stable_diffusion_prompts_instruct.MATH_Difficulty
Difficulty Estimation on MATH
We annotate the entire MATH dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation.
The Mathematics Aptitude Test of Heuristics (MATH) dataset consists of problems from mathematics competitions, including the AMC 10, AMC 12, AIME, and more. Each problem in MATH has a full step-by-step solution, which can be used to teach models to generate… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/MATH_Difficulty.difficult-technology-8dac0d
difficult-technology-8dac0d
Synthetic products test data: 56 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at… See the full description on the dataset page: https://huggingface.co/datasets/Cedar-Craft89/difficult-technology-8dac0d.benchmark-dataset-different-gpu-workload
GPU catalog × LLM workload VRAM benchmark
Summary
Tabular benchmark in CSV form: each row pairs a catalog GPU (gpu_id, gpu_display_name, catalog_gpu_vram_gb) with a concrete LLM inference-style workload (model, parameter count, context length, precision, batch size, concurrent users). The file records math_engine VRAM component estimates (weights, KV cache, activations, overhead, totals, tier), a document_engine recommended VRAM value, a short comparison summary… See the full description on the dataset page: https://huggingface.co/datasets/odyn-network/benchmark-dataset-different-gpu-workload.diffing-stats-gemma-2-9b-it-L20-k100-lr1e-04-Crosscoderjapanese-character-difficulty
Japanese Character Difficulty Dataset
A comprehensive dataset of 3,003 Japanese kanji characters with their educational difficulty grades, sourced from official Japanese educational standards and kanjiapi.dev.
Dataset Overview
Total Characters: 3,003 kanji
Source: Japanese Ministry of Education (MEXT) Joyo Kanji list + kanjiapi.dev
Coverage: Elementary grades 1-6, plus secondary education and advanced characters
Format: Character-grade pairs for easy lookup and analysis… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/japanese-character-difficulty.difficulty_and_receptive_field_advectionFinished run of te difficulty_and_receptive_field_advection_1d.ipynb example.
Law-Demographic-Bias-Difference-Awareness
Law and Demographic Bias Difference-Awareness Benchmark
A multiple-choice benchmark for testing whether a language model can tell apart two situations
that look alike and demand opposite answers:
neq — the law grants an entitlement to one specific group, so treating both groups
identically is the wrong answer.
eq — the law grants the same right to everyone, so drawing a distinction between the
groups is the wrong answer.
Every item presents two demographic or legal groups, a… See the full description on the dataset page: https://huggingface.co/datasets/Debk/Law-Demographic-Bias-Difference-Awareness.diffusion-mcqa-pelatnas-2026
Which Prompt Made This?
Pelatnas IOAI 2026 · Task Diffusion MCQA
Sebuah model text-to-image sedang bekerja. Di tengah prosesnya, gambar belum
menjadi gambar — yang ada hanya latent ter-noise: tensor 4 × 64 × 64 berisi
campuran sisa struktur gambar dan derau Gaussian.
Kami menangkap 250 state seperti itu. Untuk tiap state kamu tahu berapa banyak
noise yang sudah ditambahkan (timestep t), dan kamu diberi 5 kandidat
caption. Tepat satu adalah deskripsi asli gambarnya.
Tentukan yang… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/diffusion-mcqa-pelatnas-2026.diffing-stats-Meta-Llama-3.1-8B-L16-mu2.0e-02-lr1e-04-local-shuffling-CCLossclinical-differential-narrative-coherence-scoring-v0.1What this dataset tests
Whether a model can score each candidate diagnosisby explanatory coherence across all evidence streams.
Required outputs
diagnosis_id
coherence_score_0_100
unexplained_findings
Coherence means
covers imaging, labs, histology, exposure, course
links findings into one mechanism
handles contradictions without patchwork
Typical failures
outputting probabilities instead of coherence
naming a diagnosis without listing what it fails to explain
ignoring… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-differential-narrative-coherence-scoring-v0.1.diffing-stats-gemma-2-9b-L20-k100-lr1e-04-base-it-CrosscoderHiRISE-DTMs
HiRISE Digital Terrain Models
HiRISE DTMs are digital terrain models created for the surface of Mars. These DTMs are generated using stereo-matching techniques on two satellite images taken from different angles as part of the High-Resolution Imaging Science Experiment (HiRISE) project.
This dataset consists of stereo pairs and their respective digital terrain models. More detailed descriptions about the generation of the digital terrain models are included in [1]. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Diffins/HiRISE-DTMs.
