datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aimo-validation-aime
Dataset Card for AIMO Validation AIME
All 90 problems come from AIME 22, AIME 23, and AIME 24, and have been extracted directly from the AOPS wiki page https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions
This dataset serves as an internal validation set during our participation in the AIMO progress prize competition. Using data after 2021 is to avoid potential overlap with the MATH training set.
Here are the different columns in the dataset:
problem: the… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/aimo-validation-aime.quantum-like-attention-framework-1.3b-untuned-validation
Quantum Like Attention Framework (Q.L.A.F) 1.3b untuned
This repository contains the model checkpoints, downstream evaluation scores, and pretraining convergence logs for the Quantum Like Attention Framework (Q.L.A.F) 1.3B configuration.
Key Specifications & Architecture
Model Name: Q.L.A.F 1.3b untuned (Quantum Like Attention Framework - Hybrid Architecture)
Parameters: 1.3B parameters total configuration (327M active parameter student subset)
Layer Count: 12… See the full description on the dataset page: https://huggingface.co/datasets/IgnisCogitationis/quantum-like-attention-framework-1.3b-untuned-validation.aimo-validation-amc
Dataset Card for AIMO Validation AMC
All 83 come from AMC12 2022, AMC12 2023, and have been extracted from the AOPS wiki page https://artofproblemsolving.com/wiki/index.php/AMC_12_Problems_and_Solutions
This dataset serves as an internal validation set during our participation in the AIMO progress prize competition. Using data after 2021 is to avoid potential overlap with the MATH training set.
Here are the different columns in the dataset:
problem: the modified problem statement… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/aimo-validation-amc.esb-datasets-earnings22-validation-tiny-filteredA filtered (<=30s duration) slice (512 samples) of the Earnings22 dataset.
def add_duration(sample):
y, sr = sample['audio']["array"], sample['audio']["sampling_rate"]
sample['duration_ms']=librosa.get_duration(y=y, sr=sr) * 1000
return sample
tedlium = load_dataset("esb/datasets", "earnings22", split='validation', trust_remote_code=True)
# compute duration to filter
tedlium = tedlium.map(add_duration)
tedlium = tedlium.select(range(512))
# Whisper max supported duration
tedlium… See the full description on the dataset page: https://huggingface.co/datasets/D4nt3/esb-datasets-earnings22-validation-tiny-filtered.Imagenet-1k_validationaimo-validation-math-level-5
Dataset Card for AIMO Validation MATH Level 5
A subset of level 5 problems from https://huggingface.co/datasets/lighteval/MATH
We have extracted the final answer from boxed, and only keep those with integer outputs.
gaia_validationaimo-validation-math-level-4
Dataset Card for AIMO Validation MATH Level 4
A subset of level 4 problems from https://huggingface.co/datasets/lighteval/MATH
We have extracted the final answer from boxed, and only keep those with integer outputs.
Surya-1.0_validation_data
Validation data for Surya 1.0
This dataset comprises imagery from NASA's Solar Dynamics Observatory (SDO). The data can and should be used to validate a local installation of the Surya Foundation Model for Heliophysics. The data is compressed; you should use the hdf5plugin to read it directly.
QuantiPhy-validation
QuantiPhy (Validation Set)
Dataset Summary
QuantiPhy is a benchmark for evaluating whether vision–language models (VLMs) can perform quantitative physical inference from visual evidence, rather than producing plausible but ungrounded numerical guesses.
This repository contains the official validation set of QuantiPhy, released to support model development, ablation studies, and preliminary evaluation.The validation set represents approximately 4% of the full benchmark and… See the full description on the dataset page: https://huggingface.co/datasets/PaulineLi/QuantiPhy-validation.gaia-validation-sampled_50olmo-2-pretrain-validationxsum_validation_t5ii-agent_gaia-benchmark_validationCOCO_captions_validation
Dataset Card for "COCO_captions_validation"
More Information needed
sae-skeskinen-TinyStories-hf-validation-tokenizer-gpt2_playnq_open-validation
Dataset Card for "nq_open-validation"
More Information needed
VQAv2_validation
Dataset Card for "VQAv2_validation"
More Information needed
getting-started-labeled-validation
Dataset Card for validation_photos
This is a FiftyOne dataset with 143 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("TheSteve0/getting-started-labeled-validation")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/getting-started-labeled-validation.FineFineWeb-validation
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-validation.Places365-Validationverl_validation_mmmu_charxiv_mathversekorea_speech_mfa_aligned_validationbrand-spectrometer-validation
Brand Spectrometer — Validation Study
Reproducible validation data for the Brand Spectrometer, an instrument that reads
cohort-resolved, eight-dimensional brand-perception specifications from public artifacts
via cross-operator LLM pipelines.
This dataset accompanies the Brand Spectrometer methods paper and holds the raw,
fully-reproducible outputs of its validation battery. The instrument is ground-truth
absent by design: it does not recover a "true" brand spec, and cohort… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/brand-spectrometer-validation.seekr-world-habitat-gs-validation
Seekr World — Habitat-GS Validation Scenes
This dataset packages the nine InteriorGS validation scenes used to develop and evaluate Seekr World.
The package preserves the original Habitat navigation assets while adding browser-compatible scene metadata and navigation geometry for use with Seekr World.
It also includes a frozen Habitat-GS validation benchmark subset for ObjectNav and Vision-Language Navigation (VLN) on these nine scenes.
Dataset contents
Each scene… See the full description on the dataset page: https://huggingface.co/datasets/flavspa/seekr-world-habitat-gs-validation.COVID-QA-unique-context-test-10-percent-validation-10-percent
Dataset Card for "COVID-QA-unique-context-test-10-percent-validation-10-percent"
More Information needed
wmt14_de-en_validation_t5VQAv2_sample_validation
Dataset Card for "VQAv2_sample_validation"
More Information needed
getting-started-validation-clip-pred
Dataset Card for labeled_validation_predicted_clip
This is a FiftyOne dataset with 143 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("TheSteve0/getting-started-validation-clip-pred")
# Launch the App
session =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/getting-started-validation-clip-pred.Imagenet1k_sample_validation
Dataset Card for "Imagenet1k_sample_validation"
More Information needed
