datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sweden_100K_difficultdiffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04 Contains maximum activating examples for all the features of our crosscoder trained on gemma 2 2B layer 13 available here: https://huggingface.co/Butanium/gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04/blob/main/README.md
base_examples.pt contains all the maximum examples of the feature on a subset of validation test of fineweb
chat_examples.pt is the same but for lmsys chat data
chat_base_examples.pt is a merge of the two above files.
All files are of the type dict[int, list[tuple[float… See the full description on the dataset page: https://huggingface.co/datasets/science-of-finetuning/diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04.diffusion-mcqa-gen-pelatnas-2026
Which Prompt Made This? — Generated Edition
Pelatnas IOAI 2026 · Task Diffusion MCQA (varian trajectory)
Sebuah model text-to-image sedang bekerja. Di tengah prosesnya, gambar belum
menjadi gambar — yang ada hanya latent ter-noise: tensor 4 × 64 × 64 berisi
campuran struktur yang mulai muncul dan derau Gaussian.
Kali ini kalimat itu harfiah. Latent yang kamu terima benar-benar diambil dari
tengah proses generate: sebuah trajectory denoising DDIM 50 langkah dihentikan
sejenak… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/diffusion-mcqa-gen-pelatnas-2026.benchmarks
Welcome to 🤗 Diffusers Benchmarks!
This is dataset where we keep track of the inference latency and memory information of the core models in the diffusers library.
Currently, the core models are:
Flux
Wan
LTX
SDXL
Note that we will continue to extend this list based on their usage.
You can analyze the results in this demo.
[!IMPORTANT]
Instead of benchmarking the entire diffusion pipelines, we only benchmark the forward passes
of the diffusion networks under different settings… See the full description on the dataset page: https://huggingface.co/datasets/diffusers/benchmarks.DeepScaleR_Difficulty
Difficulty Estimation on DeepScaleR
We annotate the entire DeepScaleR dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation.
DeepScaleR is a curated dataset of 40,000 reasoning-intensive problems used to train and evaluate reinforcement learning-based methods for large language models.
Difficulty Scoring Method
Difficulty scores are estimated using the… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/DeepScaleR_Difficulty.Diffusion-Reward-Modeling-for-Text-Rendering-Dataset
🖼️ Text-to-Image Rendering Dataset
A dataset of 14k text prompts for image generation with text rendering evaluation
📚 Dataset Overview
This dataset contains 14,000 text prompts specifically designed for:
Image generation with text rendering
Evaluating text preservation in generated images
Training diffusion models for better text rendering
Each prompt comes with:
Pre-extracted target text for rendering
5 Stable Diffusion 3 generated latents (70k total)
Dual… See the full description on the dataset page: https://huggingface.co/datasets/leffff/Diffusion-Reward-Modeling-for-Text-Rendering-Dataset.diffusion-vs-ar-hard-sudoku
Diffusion vs AR Hard Sudoku
This repository packages 8,148,696 Sudoku examples in the CSV format expected
by HKUNLP/diffusion-vs-ar, plus its original 100k/1k easy baseline.
Every processed file has these columns:
column
meaning
quizzes
81 row-major digits; 0 is an empty cell
solutions
complete 81-digit solution
source
original collection
dataset
normalized dataset family
official_rating
rating supplied by the source
rating_type
semantics of that rating… See the full description on the dataset page: https://huggingface.co/datasets/fhyfhy/diffusion-vs-ar-hard-sudoku.GSM8K_Difficulty
Difficulty Estimation on DeepScaleR
We annotate the entire GSM8K dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation.
GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/GSM8K_Difficulty.orz_math_difficulty
Difficulty Estimation on Open Reasoner Zero
We annotate the entire Open Reasoner Zero dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction.
Open Reasoner Zero is a curated a dataset of 57,000 reasoning-intensive problems used to train and evaluate reinforcement learning-based methods for large language models.
Difficulty Scoring Method
Difficulty scores are estimated using… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/orz_math_difficulty.MATH_Difficulty
Difficulty Estimation on MATH
We annotate the entire MATH dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation.
The Mathematics Aptitude Test of Heuristics (MATH) dataset consists of problems from mathematics competitions, including the AMC 10, AMC 12, AIME, and more. Each problem in MATH has a full step-by-step solution, which can be used to teach models to generate… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/MATH_Difficulty.difficult-technology-8dac0d
difficult-technology-8dac0d
Synthetic products test data: 56 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at… See the full description on the dataset page: https://huggingface.co/datasets/Cedar-Craft89/difficult-technology-8dac0d.benchmark-dataset-different-gpu-workload
GPU catalog × LLM workload VRAM benchmark
Summary
Tabular benchmark in CSV form: each row pairs a catalog GPU (gpu_id, gpu_display_name, catalog_gpu_vram_gb) with a concrete LLM inference-style workload (model, parameter count, context length, precision, batch size, concurrent users). The file records math_engine VRAM component estimates (weights, KV cache, activations, overhead, totals, tier), a document_engine recommended VRAM value, a short comparison summary… See the full description on the dataset page: https://huggingface.co/datasets/odyn-network/benchmark-dataset-different-gpu-workload.diffing-stats-gemma-2-9b-it-L20-k100-lr1e-04-Crosscoderdifficulty_and_receptive_field_advectionFinished run of te difficulty_and_receptive_field_advection_1d.ipynb example.
diffusion-mcqa-pelatnas-2026
Which Prompt Made This?
Pelatnas IOAI 2026 · Task Diffusion MCQA
Sebuah model text-to-image sedang bekerja. Di tengah prosesnya, gambar belum
menjadi gambar — yang ada hanya latent ter-noise: tensor 4 × 64 × 64 berisi
campuran sisa struktur gambar dan derau Gaussian.
Kami menangkap 250 state seperti itu. Untuk tiap state kamu tahu berapa banyak
noise yang sudah ditambahkan (timestep t), dan kamu diberi 5 kandidat
caption. Tepat satu adalah deskripsi asli gambarnya.
Tentukan yang… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/diffusion-mcqa-pelatnas-2026.diffing-stats-Meta-Llama-3.1-8B-L16-mu2.0e-02-lr1e-04-local-shuffling-CCLossdiffing-stats-gemma-2-9b-L20-k100-lr1e-04-base-it-Crosscoderdiffing-stats-SAE-difference_cb-gemma-2-2b-L13-k100-x8-lr1e-04-local-shufflingdiffing-stats-SAE-base-gemma-2-2b-L13-k100-x32-lr1e-04-local-shufflingdiffing-stats-SAE-difference_cb-gemma-2-2b-L13-k100-lr1e-04-local-shufflingdifferences-in-cumulative-influenza-vaccination-co
Differences in Cumulative Influenza Vaccination Coverage by Selected Demographics, Adults 18 Year and Older, NIS Adult COVID Module
Description
Differences in Cumulative Influenza Vaccination Coverage by Selected Demographics, Adults 18 Year and Older, NIS Adult COVID Module
• The National Immunization Survey-Adult COVID Module (NIS-ACM) was launched in April 2021 among adults 18 years and older. The survey was used to monitor COVID-19 vaccination uptake and confidence in… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/differences-in-cumulative-influenza-vaccination-co.diffing-stats-gemma-2-9b-it-DPO-L20-k100-lr1e-04-dpo-simpo-Crosscoderdiffing-stats-SAE-difference_cb-gemma-2-2b-L13-k100-x2-lr1e-04-local-shufflingdiffing-stats-SAE-chat-gemma-2-2b-L13-k100-lr1e-04-local-shufflingdiffing-stats-gemma-2-2b-L13-k100-lr1e-04-local-shuffling-CCLossdiffing-stats-gemma-2-9b-it-L20-mu1.0e-01-lr1e-04-local-shuffling-CrosscoderLossdiffusers-quantization-benchmarksdiffing-stats-gemma-2-9b-L20-k100-lr1e-04-Crosscoderdiffing-stats-SAE-difference-gemma-2-2b-L13-k100-lr1e-04-local-shufflingdiffing-stats-gemma-2-9b-L20-k100-lr1e-04-base-dpo-Crosscoder
