datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aimo-validation-aime
Dataset Card for AIMO Validation AIME
All 90 problems come from AIME 22, AIME 23, and AIME 24, and have been extracted directly from the AOPS wiki page https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions
This dataset serves as an internal validation set during our participation in the AIMO progress prize competition. Using data after 2021 is to avoid potential overlap with the MATH training set.
Here are the different columns in the dataset:
problem: the… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/aimo-validation-aime.aimo-validation-amc
Dataset Card for AIMO Validation AMC
All 83 come from AMC12 2022, AMC12 2023, and have been extracted from the AOPS wiki page https://artofproblemsolving.com/wiki/index.php/AMC_12_Problems_and_Solutions
This dataset serves as an internal validation set during our participation in the AIMO progress prize competition. Using data after 2021 is to avoid potential overlap with the MATH training set.
Here are the different columns in the dataset:
problem: the modified problem statement… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/aimo-validation-amc.esb-datasets-earnings22-validation-tiny-filteredA filtered (<=30s duration) slice (512 samples) of the Earnings22 dataset.
def add_duration(sample):
y, sr = sample['audio']["array"], sample['audio']["sampling_rate"]
sample['duration_ms']=librosa.get_duration(y=y, sr=sr) * 1000
return sample
tedlium = load_dataset("esb/datasets", "earnings22", split='validation', trust_remote_code=True)
# compute duration to filter
tedlium = tedlium.map(add_duration)
tedlium = tedlium.select(range(512))
# Whisper max supported duration
tedlium… See the full description on the dataset page: https://huggingface.co/datasets/D4nt3/esb-datasets-earnings22-validation-tiny-filtered.aimo-validation-math-level-5
Dataset Card for AIMO Validation MATH Level 5
A subset of level 5 problems from https://huggingface.co/datasets/lighteval/MATH
We have extracted the final answer from boxed, and only keep those with integer outputs.
gaia_validationaimo-validation-math-level-4
Dataset Card for AIMO Validation MATH Level 4
A subset of level 4 problems from https://huggingface.co/datasets/lighteval/MATH
We have extracted the final answer from boxed, and only keep those with integer outputs.
QuantiPhy-validation
QuantiPhy (Validation Set)
Dataset Summary
QuantiPhy is a benchmark for evaluating whether vision–language models (VLMs) can perform quantitative physical inference from visual evidence, rather than producing plausible but ungrounded numerical guesses.
This repository contains the official validation set of QuantiPhy, released to support model development, ablation studies, and preliminary evaluation.The validation set represents approximately 4% of the full benchmark and… See the full description on the dataset page: https://huggingface.co/datasets/PaulineLi/QuantiPhy-validation.gaia-validation-sampled_50olmo-2-pretrain-validationii-agent_gaia-benchmark_validationCOCO_captions_validation
Dataset Card for "COCO_captions_validation"
More Information needed
sae-skeskinen-TinyStories-hf-validation-tokenizer-gpt2_playnq_open-validation
Dataset Card for "nq_open-validation"
More Information needed
VQAv2_validation
Dataset Card for "VQAv2_validation"
More Information needed
korea_speech_mfa_aligned_validationverl_validation_mmmu_charxiv_mathverseVQAv2_sample_validation
Dataset Card for "VQAv2_sample_validation"
More Information needed
COVID-QA-unique-context-test-10-percent-validation-10-percent
Dataset Card for "COVID-QA-unique-context-test-10-percent-validation-10-percent"
More Information needed
Imagenet1k_sample_validation
Dataset Card for "Imagenet1k_sample_validation"
More Information needed
TextVQA_validation
Dataset Card for "TextVQA_validation"
More Information needed
Korea-AIHub-middlesenior-dialect-speech-validation-part2pile-validationThis dataset has been created as an artefact of the paper Causal Estimation of Memorisation Profiles (Lesci et al., 2024).
More info about this dataset in the related collection Memorisation-Profiles.
The validation data used in our study. The Pythia suite does not have an official validation. However, we confirmed with the authors that the Pile validation split (this one) was not seen during training.
It is still a bit confusing whether the Pile data can be released freely. Thus, we will… See the full description on the dataset page: https://huggingface.co/datasets/pietrolesci/pile-validation.CIVQA_EasyOCR_Validation
CIVQA EasyOCR Validation Dataset
The CIVQA (Czech Invoice Visual Question Answering) dataset was created with EasyOCR. This dataset contains only the validation split. The train part of the dataset can be found on this URL: https://huggingface.co/datasets/fimu-docproc-research/CIVQA_EasyOCR_Train
The encoded validation dataset for the LayoutLM can be found on this link: https://huggingface.co/datasets/fimu-docproc-research/CIVQA_EasyOCR_LayoutLM_Validation
All invoices used in this… See the full description on the dataset page: https://huggingface.co/datasets/fimu-docproc-research/CIVQA_EasyOCR_Validation.wikipedia-20220301.en-0.005-validationVPPO_MMK12_validation
Dataset Card for VPPO_MMK12_validation
Dataset Details
Dataset Description
This dataset is the official validation split used to fine-tune the VPPO-7B and VPPO-32B models presented in our paper, "Spotlight on Token Perception for Multimodal Reinforcement Learning".
This is a direct copy of the test split of FanqingM/MMK12 dataset. We have isolated it here to ensure the exact version used in our experiments is publicly available, guaranteeing reproducibility for… See the full description on the dataset page: https://huggingface.co/datasets/chamber111/VPPO_MMK12_validation.c4-en-validationdapo_validationttm-validation-datasetdrawvla-prompt-validation-clean
DrawVLA — Sketch-Prompt Validation
Circle (which) + arrow (where) + caption (what) visual instructions overlaid on
LIBERO observations, each labelled with a binary
verdict for training a prompt validator or a self-checking VLA:
right — every channel is correct and exactly one reading survives; execute.
wrong — a channel is incorrect or the deictic prompt remains under-determined;
reject. Formerly ambiguous prompts are retained in this class.
All captions are name-free L2/L3… See the full description on the dataset page: https://huggingface.co/datasets/shibuina/drawvla-prompt-validation-clean.VizWiz_validation
Dataset Card for "VizWiz_validation"
More Information needed
