CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AI-MO /aimo-validation-aime Dataset Card for AIMO Validation AIME All 90 problems come from AIME 22, AIME 23, and AIME 24, and have been extracted directly from the AOPS wiki page https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions This dataset serves as an internal validation set during our participation in the AIMO progress prize competition. Using data after 2021 is to avoid potential overlap with the MATH training set. Here are the different columns in the dataset: problem: the… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/aimo-validation-aime.textn<1K68 likes37k downloads1y agoHugging Face02IgnisCogitationis /quantum-like-attention-framework-1.3b-untuned-validation Quantum Like Attention Framework (Q.L.A.F) 1.3b untuned This repository contains the model checkpoints, downstream evaluation scores, and pretraining convergence logs for the Quantum Like Attention Framework (Q.L.A.F) 1.3B configuration. Key Specifications & Architecture Model Name: Q.L.A.F 1.3b untuned (Quantum Like Attention Framework - Hybrid Architecture) Parameters: 1.3B parameters total configuration (327M active parameter student subset) Layer Count: 12… See the full description on the dataset page: https://huggingface.co/datasets/IgnisCogitationis/quantum-like-attention-framework-1.3b-untuned-validation.3 likes27k downloads5m agoHugging Face03AI-MO /aimo-validation-amc Dataset Card for AIMO Validation AMC All 83 come from AMC12 2022, AMC12 2023, and have been extracted from the AOPS wiki page https://artofproblemsolving.com/wiki/index.php/AMC_12_Problems_and_Solutions This dataset serves as an internal validation set during our participation in the AIMO progress prize competition. Using data after 2021 is to avoid potential overlap with the MATH training set. Here are the different columns in the dataset: problem: the modified problem statement… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/aimo-validation-amc.tabularn<1K19 likes11k downloads1y agoHugging Face04D4nt3 /esb-datasets-earnings22-validation-tiny-filteredA filtered (<=30s duration) slice (512 samples) of the Earnings22 dataset. def add_duration(sample): y, sr = sample['audio']["array"], sample['audio']["sampling_rate"] sample['duration_ms']=librosa.get_duration(y=y, sr=sr) * 1000 return sample tedlium = load_dataset("esb/datasets", "earnings22", split='validation', trust_remote_code=True) # compute duration to filter tedlium = tedlium.map(add_duration) tedlium = tedlium.select(range(512)) # Whisper max supported duration tedlium… See the full description on the dataset page: https://huggingface.co/datasets/D4nt3/esb-datasets-earnings22-validation-tiny-filtered.audion<1K0 likes7.9k downloads2y agoHugging Face05Tsomaros /Imagenet-1k_validationimage10K<n<100K0 likes3.3k downloads2y agoHugging Face06AI-MO /aimo-validation-math-level-5 Dataset Card for AIMO Validation MATH Level 5 A subset of level 5 problems from https://huggingface.co/datasets/lighteval/MATH We have extracted the final answer from boxed, and only keep those with integer outputs. textn<1K11 likes2.5k downloads2y agoHugging Face07jiyu9437 /gaia_validationtextn<1K0 likes1.5k downloads1y agoHugging Face08AI-MO /aimo-validation-math-level-4 Dataset Card for AIMO Validation MATH Level 4 A subset of level 4 problems from https://huggingface.co/datasets/lighteval/MATH We have extracted the final answer from boxed, and only keep those with integer outputs. textn<1K4 likes1.3k downloads2y agoHugging Face09nasa-ibm-ai4science /Surya-1.0_validation_data Validation data for Surya 1.0 This dataset comprises imagery from NASA's Solar Dynamics Observatory (SDO). The data can and should be used to validate a local installation of the Surya Foundation Model for Heliophysics. The data is compressed; you should use the hdf5plugin to read it directly. 1 likes1.2k downloads1y agoHugging Face10PaulineLi /QuantiPhy-validation QuantiPhy (Validation Set) Dataset Summary QuantiPhy is a benchmark for evaluating whether vision–language models (VLMs) can perform quantitative physical inference from visual evidence, rather than producing plausible but ungrounded numerical guesses. This repository contains the official validation set of QuantiPhy, released to support model development, ablation studies, and preliminary evaluation.The validation set represents approximately 4% of the full benchmark and… See the full description on the dataset page: https://huggingface.co/datasets/PaulineLi/QuantiPhy-validation.tabularvideo-text-to-textn<1K7 likes1.1k downloads9mo agoHugging Face11lauspectrum /gaia-validation-sampled_50textn<1K0 likes1k downloads1y agoHugging Face12sbordt /olmo-2-pretrain-validationtext10K<n<100K0 likes1k downloads5mo agoHugging Face13embedded-language-flows /xsum_validation_t50 likes937 downloads4mo agoHugging Face14Intelligent-Internet /ii-agent_gaia-benchmark_validationtextn<1K8 likes926 downloads1y agoHugging Face15Multimodal-Fatima /COCO_captions_validation Dataset Card for "COCO_captions_validation" More Information needed image1K<n<10K0 likes840 downloads4y agoHugging Face16apollo-research /sae-skeskinen-TinyStories-hf-validation-tokenizer-gpt2_playtext10K<n<100K0 likes741 downloads3y agoHugging Face17seonglae /nq_open-validation Dataset Card for "nq_open-validation" More Information needed text100K<n<1M0 likes569 downloads3y agoHugging Face18Multimodal-Fatima /VQAv2_validation Dataset Card for "VQAv2_validation" More Information needed image100K<n<1M0 likes529 downloads3y agoHugging Face19Voxel51 /getting-started-labeled-validation Dataset Card for validation_photos This is a FiftyOne dataset with 143 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("TheSteve0/getting-started-labeled-validation") # Launch the App session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/getting-started-labeled-validation.imageimage-classificationn<1K0 likes525 downloads2y agoHugging Face20m-a-p /FineFineWeb-validation FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-validation.tabulartext-classification10K<n<100K1 likes501 downloads2y agoHugging Face21dpdl-benchmark /Places365-Validationimage10K<n<100K0 likes478 downloads1y agoHugging Face22112bb /verl_validation_mmmu_charxiv_mathversetext1K<n<10K0 likes455 downloads8mo agoHugging Face23AdoCleanCode /korea_speech_mfa_aligned_validationaudio100K<n<1M0 likes455 downloads8mo agoHugging Face24spectralbranding /brand-spectrometer-validation Brand Spectrometer — Validation Study Reproducible validation data for the Brand Spectrometer, an instrument that reads cohort-resolved, eight-dimensional brand-perception specifications from public artifacts via cross-operator LLM pipelines. This dataset accompanies the Brand Spectrometer methods paper and holds the raw, fully-reproducible outputs of its validation battery. The instrument is ground-truth absent by design: it does not recover a "true" brand spec, and cohort… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/brand-spectrometer-validation.tabularn<1K0 likes447 downloads2mo agoHugging Face25flavspa /seekr-world-habitat-gs-validation Seekr World — Habitat-GS Validation Scenes This dataset packages the nine InteriorGS validation scenes used to develop and evaluate Seekr World. The package preserves the original Habitat navigation assets while adding browser-compatible scene metadata and navigation geometry for use with Seekr World. It also includes a frozen Habitat-GS validation benchmark subset for ObjectNav and Vision-Language Navigation (VLN) on these nine scenes. Dataset contents Each scene… See the full description on the dataset page: https://huggingface.co/datasets/flavspa/seekr-world-habitat-gs-validation.0 likes445 downloads20d agoHugging Face26minh21 /COVID-QA-unique-context-test-10-percent-validation-10-percent Dataset Card for "COVID-QA-unique-context-test-10-percent-validation-10-percent" More Information needed tabular1K<n<10K0 likes441 downloads3y agoHugging Face27embedded-language-flows /wmt14_de-en_validation_t50 likes440 downloads5mo agoHugging Face28Multimodal-Fatima /VQAv2_sample_validation Dataset Card for "VQAv2_sample_validation" More Information needed image1K<n<10K0 likes434 downloads3y agoHugging Face29Voxel51 /getting-started-validation-clip-pred Dataset Card for labeled_validation_predicted_clip This is a FiftyOne dataset with 143 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("TheSteve0/getting-started-validation-clip-pred") # Launch the App session =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/getting-started-validation-clip-pred.imageimage-classificationn<1K0 likes400 downloads2y agoHugging Face30Multimodal-Fatima /Imagenet1k_sample_validation Dataset Card for "Imagenet1k_sample_validation" More Information needed image1K<n<10K0 likes357 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.