datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
math-contests-2026
Math Contests 2026 (🔗 notadib/math-contests-2026)
197 problems from national olympiads and team-selection tests held January 2026 and onward — a held-out benchmark for math reasoning, sourced after the contests ran but before solutions were widely propagated, so they should not appear in any current LLM training data.
Excluded: any contest held in 2025 — BMO Round 1 (Nov 2025), USA TSTST, USA TST (Dec 2025) and Bundeswettbewerb Mathematik (Dec 2025) — kept strictly to events… See the full description on the dataset page: https://huggingface.co/datasets/notadib/math-contests-2026.AOBench
AOBench: Advanced Olympiad Benchmark
AOBench contains 35 hard, non-geometry problems from 2026 mathematical olympiads and selection tests, with reviewed proofs, strategy sketches of at most 25 words, and three-step outlines. Its size and difficulty are comparable to the IMO-ProofBench Advanced dataset. We also include the 22 non-geometry problems subset used in our study, with matching annotations.
The dataset accompanies When Does Long Thinking Help? Predicting Mathematical… See the full description on the dataset page: https://huggingface.co/datasets/notadib/AOBench.vga_nota_dataset
LIBERO-3D — GT depth + camera for openvla/modified_libero_rlds
Metric depth and camera pose (K, E) for every training episode of
openvla/modified_libero_rlds
(LIBERO, 4 suites, no_noops). The depth and camera are rendered from the LIBERO MuJoCo
simulator and aligned frame-for-frame to the RLDS episodes. Self-contained: the RGB the policy
sees is bundled in too — those frames are not re-rendered here; they are copied verbatim from
the source dataset openvla/modified_libero_rlds.… See the full description on the dataset page: https://huggingface.co/datasets/nota-gmbh/vga_nota_dataset.chess-notation-datasetThis dataset contains 9500+ snapshots from lichess rated games from January/February 2015. The snapshots are in FEN, and the answer column consists of a JSON representing the gameboard as described in the FEN.
Intended to be used with my ChessNotationEnvironment on Prime Environment hub, but feel free to use however you wish!
sdf-data-meta_sdf_notag_dist_negNotASI__FineTome-Llama3.2-1B-0929-details
Dataset Card for Evaluation run of NotASI/FineTome-Llama3.2-1B-0929
Dataset automatically created during the evaluation run of model NotASI/FineTome-Llama3.2-1B-0929
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NotASI__FineTome-Llama3.2-1B-0929-details.NotASI__FineTome-Llama3.2-3B-1002-details
Dataset Card for Evaluation run of NotASI/FineTome-Llama3.2-3B-1002
Dataset automatically created during the evaluation run of model NotASI/FineTome-Llama3.2-3B-1002
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NotASI__FineTome-Llama3.2-3B-1002-details.NotASI__FineTome-v1.5-Llama3.2-3B-1007-details
Dataset Card for Evaluation run of NotASI/FineTome-v1.5-Llama3.2-3B-1007
Dataset automatically created during the evaluation run of model NotASI/FineTome-v1.5-Llama3.2-3B-1007
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NotASI__FineTome-v1.5-Llama3.2-3B-1007-details.sdf-data-meta_sdf_notag_dist_possdf-data-meta_sdf_notag_prox_negsdf-data-meta_sdf_notag_prox_postest_datasetNotASI__FineTome-v1.5-Llama3.2-1B-1007-details
Dataset Card for Evaluation run of NotASI/FineTome-v1.5-Llama3.2-1B-1007
Dataset automatically created during the evaluation run of model NotASI/FineTome-v1.5-Llama3.2-1B-1007
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NotASI__FineTome-v1.5-Llama3.2-1B-1007-details.notam-typed-decisions
NOTAM typed decisions
Frozen, hashed evaluation suites for typed decisions about NOTAMs — a choice, a
yes/no, or a score, each with a confidence — plus a catalogue of how NOTAMs describe
areas in free text. Built by Airside Labs so that any model,
served any way, can be scored on the same rows and read with the same per-class table
and calibration curve.
Not for operational use. These suites and the numbers quoted here are for research
and for triage tooling. NOTAMs are… See the full description on the dataset page: https://huggingface.co/datasets/AirsideLabs/notam-typed-decisions.
