datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
competition_math
Dataset Card for Mathematics Aptitude Test of Heuristics (MATH) dataset
Dataset Summary
The Mathematics Aptitude Test of Heuristics (MATH) dataset consists of problems
from mathematics competitions, including the AMC 10, AMC 12, AIME, and more.
Each problem in MATH has a full step-by-step solution, which can be used to teach
models to generate answer derivations and explanations.
Supported Tasks and Leaderboards
[More Information Needed]
Languages… See the full description on the dataset page: https://huggingface.co/datasets/qwedsacf/competition_math.soma-competition-datasetquora
Dataset Card for "quora"
Dataset Summary
The Quora dataset is composed of question pairs, and the task is to determine if the questions are paraphrases of each other (have the same meaning).
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
default
Size of downloaded dataset files: 58.17 MB
Size of the generated dataset: 58.15 MB
Total amount… See the full description on the dataset page: https://huggingface.co/datasets/quora-competitions/quora.Explorer_LLM_Rec_Competition
Explorer_LLM_Rec_Competition
This dataset contains user historical behaviors and content metadata, constructed from real user interaction histories. It covers behavior sequences across multiple domains for a single user and supports cross-domain recommendation, semantic-ID retrieval / generation, content understanding, and related tasks.
Files
File
Description
OneReason_UserProfile/
Per-user multi-domain behavior records (~500k rows)… See the full description on the dataset page: https://huggingface.co/datasets/OpenOneRec/Explorer_LLM_Rec_Competition.competition_mathcompetition_math
Dataset Card for "competition_math"
More Information needed
competition_math_hf_dataset
Dataset Card for "competition_math_hf_dataset"
Homepage - https://huggingface.co/datasets/hendrycks/competition_math
This is just the competetion math dataset, put in HF dataset format for ease of use with any finetuning tasks
MMFMChallengeLibriBrain-Competition-2026competition_math_imagessft-ready-hendrycks-competition_mathicil-competition-results
ICIL competition results
The signed, append-only result store of the RoboTensor one-demonstration in-context imitation
learning competition. The dashboard is a pure reader of this
repository: everything a visitor sees is here, and anyone can check it without trusting the site.
What a duel publishes
Two models run an identical unit list — same tasks, same initial states, same demonstration
per unit, same seeds. Each skill is one success rate; the score is their… See the full description on the dataset page: https://huggingface.co/datasets/robotensor/icil-competition-results.competition_math
Dataset Card for "competition_math"
Added column with final solution extracted from \boxed{} tags.
Added numeric congig that only contains questions with numeric answers.
Dataset Summary
MATH contains 12,500 challenging competition mathematics problems. Each problem in MATH has a full step-by-step solution which can be used to teach models to generate answer derivations and explanation
This dataset card aims to be a base template for new datasets.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/jeggers/competition_math.korean_mmqa_competitionhighschool_math_competition初高中 中文数学竞赛pdf书籍。
deepseek_ocr.zip是使用https://github.com/deepseek-ai/DeepSeek-OCR/得到的OCR文本
RealPDE-Competition-Data
RealPDE Competition Data (NeurIPS 2026)
Training data and baseline checkpoints for the NeurIPS 2026 RealPDE
Competition. This is a mirror of the
competition's Google Drive release, hosted here because the Drive link runs into
a per-file anonymous download quota when many people fetch it at once.
Both tracks share this release:
Track 1, Sim2Real — codabench.org/competitions/17363
Track 2, LTTTA — codabench.org/competitions/17385
Contents
train_sim.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/AI4Science-WestlakeU/RealPDE-Competition-Data.forum-competition-math-training-pool
Forum competition mathematics training pool
Olympiad and contest mathematics from three public datasets, gathered at pinned revisions and
shipped twice over. sources/ holds each dataset the way its publisher ships it, in its own file
format with its own fields and nothing renamed, 287091 rows across three folders. pool/ holds the
union of those same datasets in one format, one JSON object per line, deduplicated by problem text
and reduced to 282140 rows, every row labelled with… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/forum-competition-math-training-pool.secret-loyalty-competition-data
Secret-loyalty organisms — training banks and eval batteries
The data behind KKing23/secret-loyalty-competition-organisms.
Code and full result trail: github.com/kaustubhkislay/secret-loyalty-competition.
Why this exists separately from the adapters. The adapters are reproducible from these
banks for the price of GPU time. These banks are not reproducible — they were written by
an LLM generator, so regenerating gives different data and every published number becomes… See the full description on the dataset page: https://huggingface.co/datasets/KKing23/secret-loyalty-competition-data.RO-competition-paper-datasetcompetition-math-training-pool
Competition mathematics training pool
Public competition mathematics, six datasets gathered at pinned revisions, shipped twice over.
sources/ holds each dataset the way its publisher ships it, in its own file format with its own
fields and nothing renamed, 1951046 rows across six folders. pool/ holds the union of those
same datasets in one format, one JSON object per line, deduplicated by problem text and reduced
to 1125451 rows, every row labelled with the dataset it came from… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/competition-math-training-pool.competition_math_selectedBengali_Competition_DatasetCompetition-Submissions
Competition Submissions
A curated dataset of writing that models compassionate moral reasoning about nonhuman sentient beings — animals, insects, digital minds, and entities whose moral status is uncertain.
Designed for pretraining and fine-tuning language models to reason more carefully and compassionately when facing decisions that affect sentient life.
Why This Dataset Exists
Recent alignment research shows that training on synthetic documents depicting… See the full description on the dataset page: https://huggingface.co/datasets/Hyperstition-for-Good/Competition-Submissions.GWFSS-competition
Introduction
Competition Page
If you want any update on the Global Wheat Dataset Community, go on https://www.global-wheat.com/
Wheat is a cornerstone of global food security, serving as a dietary staple for billions of people worldwide. Detailed analysis of wheat plants can help scientists and farmers cultivate healthier, more resilient, and more productive crops. The Global Wheat Full Semantic Segmentation (GWFSS) task aims to perform pixel-level segmentation of plant components… See the full description on the dataset page: https://huggingface.co/datasets/XIANG-Shuai/GWFSS-competition.adult-census-competitionMaths_competition_questionsExplorer_LLM_Rec_Competition
Explorer_LLM_Rec_Competition
This dataset contains user historical behaviors and content metadata, constructed from real user interaction histories. It covers behavior sequences across multiple domains for a single user and supports cross-domain recommendation, semantic-ID retrieval / generation, content understanding, and related tasks.
Files
File
Description
OneReason_UserProfile/
Per-user multi-domain behavior records (~500k rows)… See the full description on the dataset page: https://huggingface.co/datasets/akk666/Explorer_LLM_Rec_Competition.autoscientist-competition-datasetsOpenMathInstruct-2-augmented-mathnvidia/OpenMathInstruct-2のaugmented_mathだけ抜き出したものです.licenseは元データと同じです.
turkish-competition-authority-decisions
Turkish Competition Authority Decisions (Rekabet Kurulu Kararları), 1997–2026
The complete published decision history of the Turkish Competition Authority
(Rekabet Kurumu) — every Competition Board decision the regulator has made public,
in full text, with derived structural metadata.
10,367 decisions · 113,297 pages · 323 million characters · 29 years
Every decision carries its outcome, the articles of Law 4054 it turns on, the
panel that decided it (as stable pseudonymous ids… See the full description on the dataset page: https://huggingface.co/datasets/emirms/turkish-competition-authority-decisions.
