datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
math_stratos_scale_judged_and_annotated_with_difficultyExercise-Synthetic-split-ncert-chapter-mapped_filtered_difficulty_scoreddifficulty-E2H-AMC-generations
Generations Dataset: E2H-AMC
LLM-generated solutions across train/validation/test splits for multiple models.
Columns
Column
Type
Description
problem
str
Problem statement
generated_solutions
list
Generated solutions with scores
success_rate
float
Fraction of correct generations
majority_vote_is_correct
int (0/1)
Whether majority vote is correct
k
int
Number of samples generated
temperature
float
Sampling temperature
max_len
int
Maximum… See the full description on the dataset page: https://huggingface.co/datasets/CoffeeGitta/difficulty-E2H-AMC-generations.aozora-text-difficulty
Aozora Text Difficulty Dataset
This dataset contains Japanese literary texts from the Aozora Bunko digital library, enhanced with jReadability-based difficulty analysis for Japanese language learning and curriculum development.
Dataset Overview
Source: Aozora Bunko (青空文庫) - Japan's premier digital library of public domain literature
Enhancement: jReadability-based difficulty scoring using research-backed Japanese readability models
Primary Methodology: jReadability - A… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/aozora-text-difficulty.BlindLoop-Difficulty-Feedback
BlindLoop Difficulty Feedback
This is the public, hash-bound release of BlindLoop Section 3. Coding agents
generated executable visual-question tasks; each task's inverse program checked
the answer from rendered pixels. For complete feedback transactions, the exact
same five images were evaluated by three frontier VLMs and the resulting
difficulty signal was returned to the next generation episode.
Contents
Config
Unit
Rows
tasks
generated task
266… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/BlindLoop-Difficulty-Feedback.seed_math_exploit_difficulty_annotationdifficulty-gsm8k-generations
Generations Dataset: gsm8k
LLM-generated solutions across train/validation/test splits for multiple models.
Columns
Column
Type
Description
problem
str
Problem statement
generated_solutions
list
Generated solutions with scores
success_rate
float
Fraction of correct generations
majority_vote_is_correct
int (0/1)
Whether majority vote is correct
k
int
Number of samples generated
temperature
float
Sampling temperature
max_len
int
Maximum… See the full description on the dataset page: https://huggingface.co/datasets/CoffeeGitta/difficulty-gsm8k-generations.difficulty-aime_2025-generations
Generations Dataset: aime_2025
Paper: LLMs Encode Their Failures: Predicting Success from Pre-Generation ActivationsCode: GitHub
LLM-generated solutions across train/validation/test splits for multiple models.
Columns
Column
Type
Description
problem
str
Problem statement
generated_solutions
list
Generated solutions with scores
success_rate
float
Fraction of correct generations
majority_vote_is_correct
int (0/1)
Whether majority vote is correct
k… See the full description on the dataset page: https://huggingface.co/datasets/CoffeeGitta/difficulty-aime_2025-generations.difficulty-eval-64difficulty-MATH-generations
Generations Dataset: MATH
LLM-generated solutions across train/validation/test splits for multiple models.
Columns
Column
Type
Description
problem
str
Problem statement
generated_solutions
list
Generated solutions with scores
success_rate
float
Fraction of correct generations
majority_vote_is_correct
int (0/1)
Whether majority vote is correct
k
int
Number of samples generated
temperature
float
Sampling temperature
max_len
int
Maximum… See the full description on the dataset page: https://huggingface.co/datasets/CoffeeGitta/difficulty-MATH-generations.difficulty_filtering_seed_mathR2E-Gym-Lite-with-DifficultyKorean-DeepMath-with-Difficulty
Korean DeepMath with Difficulty
This dataset enriches ChuGyouk/Korean-DeepMath with difficulty and topic metadata from zwhe99/DeepMath-103K.
Join procedure
Rows are matched using Korean-DeepMath[extra_info][index] -> original DeepMath row index.
Added fields
original_index
difficulty
topic
Intended use
Prepared for controlled Korean mathematical reasoning SFT experiments, including difficulty-aware sampling such as TDCS.
No Easy /… See the full description on the dataset page: https://huggingface.co/datasets/Seungjun/Korean-DeepMath-with-Difficulty.japanese-math-empirical-difficulty-pilot-50k
Japanese Math Empirical Difficulty Pilot 50k
This dataset is a 50,000-problem empirical difficulty pilot, not a full empirical labeling of the original 5.66M-row source dataset.
It was created for LLM-jp experiment 0399, Team Victory SFT, to validate empirical difficulty label distribution, downstream split behavior, and the rollout/scoring pipeline before attempting labeling at the full 5.6M scale.
Current Status
This upload uses the v3 scorer with assistant-only… See the full description on the dataset page: https://huggingface.co/datasets/argo11/japanese-math-empirical-difficulty-pilot-50k.R2E-Gym-Lite-Difficulty-Clean-Onlyjapanese-text-difficulty
Aozora Text Difficulty Dataset
This dataset contains Japanese literary texts from the Aozora Bunko digital library, enhanced with jReadability-based difficulty analysis for Japanese language learning and curriculum development.
Dataset Overview
Source: Aozora Bunko (青空文庫) - Japan's premier digital library of public domain literature
Enhancement: jReadability-based difficulty scoring using research-backed Japanese readability models
Primary Methodology: jReadability - A… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/japanese-text-difficulty.numinamath_verifiable_cleaned_wo_geo_mc_difficultydifficulty_sorting_easy_seed_math_w_openthoughtsCountdown-Tasks-Difficulty-Linear-Ratio-1.2kjapanese-text-difficulty-2leveltool-use-adaptation-difficulty
Tool-Use Adaptation — Task Difficulty (gemini & Qwen)
Empirical difficulty annotations for tool-use / agentic tasks, part of an "adapting to a new tool" RL domain.
Difficulty is gauged by running two strong solvers N=32 times per task and scoring each rollout against the
task's local programmatic gold (execution / state-check / exact-args), then aggregating.
Solvers
gemini-3-5-flash-fair (gemini_* columns)
Qwen3.6-35B-A3B (qwen_* columns)
Difficulty… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/tool-use-adaptation-difficulty.MathInstruct-Core-DifficultyAware
Dataset Card for "MathInstruct-Core-DifficultyAware"
More Information needed
secondfiltered-math220k-difficulty_stratified_10k_filtered_only_medium_difficulty
Dataset Details
This dataset origins from Blancy/secondfiltered-math220k-difficulty_stratified_10k, the orginal index means the index inBlancy/secondfiltered-math220k-difficulty_stratified_10k.
secondfiltered-math220k-difficulty_stratified_10k_tokenknowndifficulty_sorting_high_seed_math_w_openthoughtsCountdown-Tasks-Difficulty-1.2kdifficulty-aime_1983-2024-generations
Generations Dataset: aime_1983-2024
LLM-generated solutions across train/validation/test splits for multiple models.
Columns
Column
Type
Description
problem
str
Problem statement
generated_solutions
list
Generated solutions with scores
success_rate
float
Fraction of correct generations
majority_vote_is_correct
int (0/1)
Whether majority vote is correct
k
int
Number of samples generated
temperature
float
Sampling temperature
max_len
int… See the full description on the dataset page: https://huggingface.co/datasets/CoffeeGitta/difficulty-aime_1983-2024-generations.DeepMath-0.5K-filteredv2-difficultydifficulty_tagauto-problems-20250318-remain_difficulty
