datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
verifiable-coding-problems-python
Dataset Card for Verifiable Coding Problems Python 10k
This dataset contains all Python problems from PrimeIntellect's verifiable-coding-problems dataset. We have formatted the verification_info and metadata columns to be proper dictionaries, but otherwise the data is the same. Please see their dataset for more details.
verifiable-coding-problems-python_decontaminated-testedverifiable-coding-problems-python_decontaminatedverifiable-coding-problems-python_decontaminated-tested-shuffledtaocp_open_problems
TAOCP Open Problems
Collection of open research problems singled out by Donald Knuth in
The Art of Computer Programming series. Its main purpose is to help measure how frontier models understand, investigate, and make verifiable progress on hard but interesting open problems.
Contents
The dataset contains 9 exercises rated 50, M50, or HM50 in the six
TAOCP editions and draft bundles available to this project. Knuth uses these
ratings for problems that were not… See the full description on the dataset page: https://huggingface.co/datasets/sytelus/taocp_open_problems.EternalMath-open-problems
EternalMath Open Problems
This dataset is the Hugging Face viewer-friendly release of the open companion problem set for EternalMath. It contains 6,049 parameterized math problems across four batches.
Batches
Batch
Rows
Language
QC status
20260325
988
English
QC-passed
anon1
1,640
Chinese
Unfiltered
anon2
1,341
Chinese
Unfiltered
anon3
2,080
Chinese
Unfiltered
Files
The viewer loads the Parquet shards in data/ as a single train… See the full description on the dataset page: https://huggingface.co/datasets/shhendu/EternalMath-open-problems.open-code-reasoning-rlvr-original-problems
