CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /ai2_arc Dataset Card for "ai2_arc" Dataset Summary A new dataset of 7,787 genuine grade-school level, multiple-choice science questions, assembled to encourage research in advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. We are also including a corpus of over 14 million science sentences… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ai2_arc.textquestion-answering1K<n<10K402 likes871k downloads3y agoHugging Face02applied-ai-018 /pretraining_v1-omega_bookstabular100M<n<1B25 likes447k downloads2y agoHugging Face03applied-ai-018 /peacock-data-public-datasets-idc0 likes376k downloads2y agoHugging Face04AI-MO /NuminaMath-CoT Dataset Card for NuminaMath CoT Dataset Summary Approximately 860k math problems, where each solution is formatted in a Chain of Thought (CoT) manner. The sources of the dataset range from Chinese high school math exercises to US and international mathematics olympiad competition problems. The data were primarily collected from online exam paper PDFs and mathematics discussion forums. The processing steps include (a) OCR from the original PDFs, (b) segmentation… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/NuminaMath-CoT.texttext-generation100K<n<1M603 likes224k downloads2y agoHugging Face05AiEDA /iDATA Dataset A dataset of AI + EDA iDATA is a dataset of AI + EDA, which can be used to train AI models for design PPA prediction, PPA-aware physical design, and related tasks. Dataset structure Describe the dataset structure. aes/ ├── iEDA_route_process_data/ # Process data exported by iEDA-iRT 2D routing ├── syn_netlist/ # The synthesized netlist files、sdc files ├── place/ # The place stage def、sdc、vectors └── route/ # The route… See the full description on the dataset page: https://huggingface.co/datasets/AiEDA/iDATA.8 likes211k downloads9mo agoHugging Face06fixie-ai /common_voice_17_0audio10M<n<100M18 likes194k downloads2y agoHugging Face07AI-MO /olympiads AI-MO Olympiad Reference Dataset This dataset contains a structured collection of Olympiad problems and their solutions, organized by competition. Contains high quality data, prioritizing "official" solutions to problems. Structure <competition name>/ # Problems and solutions from the International Mathematical Olympiad ├── raw/ # Raw problem/solution statements (.pdf) │ ├── file1.pdf │ ├── file2.pdf ├── download_script/ # the scripts used to… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/olympiads.document10 likes173k downloads11mo agoHugging Face08markov-ai /cad-1000-hours CAD-1K Open v2 - 1,018.1229 Hours 509 end-to-end, single-display Windows CAD task recordings across seven CAD software families. Each task contains: task_desc.json - task prompt, application, reference-input paths, and expected deliverables input_files/ - reference inputs named input.ext or input_N.ext output_files/ - submitted CAD deliverables and supplemental outputs named output.ext or output_N.ext rubrics.json - task-specific evaluation criteria task_overview.pdf - review… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/cad-1000-hours.document498 likes151k downloads5d agoHugging Face09SakanaAI /AI-CUDA-Engineer-Archive The AI CUDA Engineer Archive 👷: Agentic CUDA Kernel Discovery, Optimization & Composition We release The AI CUDA Engineer archive, a dataset consisting of approximately 30,000 CUDA kernels generated by The AI CUDA Engineer. It is released under the CC-By-4.0 license and can be accessed via HuggingFace and interactively visualized here. The dataset is based on the Kernel tasks provided in KernelBench and includes a torch reference implementation, torch, NCU and Clang-tidy… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/AI-CUDA-Engineer-Archive.tabular10K<n<100K226 likes136k downloads2y agoHugging Face10Maxwell-Jia /AIME_2024 AIME 2024 Dataset Dataset Description This dataset contains problems from the American Invitational Mathematics Examination (AIME) 2024. AIME is a prestigious high school mathematics competition known for its challenging mathematical problems. Dataset Details Format: JSONL Size: 30 records Source: AIME 2024 I & II Language: English Data Fields Each record contains the following fields: ID: Problem identifier (e.g., "2024-I-1" represents Problem 1… See the full description on the dataset page: https://huggingface.co/datasets/Maxwell-Jia/AIME_2024.texttext-generationn<1K86 likes101k downloads2y agoHugging Face11ropedia-ai /xperience-10mgated ⚠️ Important: If you have already submitted an access request but have not completed the required DocuSign agreement, your request will remain pending. Please complete signing and we will grant access once verified. Interactive Intelligence from Human Xperience Xperience-10M Dataset Summary Xperience-10M is a large-scale egocentric multimodal dataset of human experience for embodied AI, robotics, world models, and spatial… See the full description on the dataset page: https://huggingface.co/datasets/ropedia-ai/xperience-10m.3dvideo-classification1M<n<10M249 likes99k downloads5mo agoHugging Face12math-ai /aime25 AIME 25 American Invitational Mathematics Examination (AIME) 2025 Citation If you use the AIME25 dataset in your research, please consider citing it as follows: @misc{aime25, title={American Invitational Mathematics Examination (AIME) 2025}, author={Zhang, Yifan and Math-AI, Team}, year={2025}, } textn<1K38 likes96k downloads8mo agoHugging Face13aisa-group /PostTrainBench-Trajectories PostTrainBench Agent Traces Agent traces from PostTrainBench (GitHub), a benchmark that measures CLI agents' ability to post-train base LLMs. Task Each agent is given: A pre-trained base LLM to fine-tune An evaluation script for a specific benchmark 10 hours on an NVIDIA H100 80GB GPU The agent must autonomously improve the model's performance on the target benchmark using any post-training strategy it chooses (SFT, LoRA, RLHF, prompt engineering for data… See the full description on the dataset page: https://huggingface.co/datasets/aisa-group/PostTrainBench-Trajectories.text-generationn<1K8 likes84k downloads1d agoHugging Face14yaak-ai /L2DTL;DR of L2D, the world's largest self-driving dataset! Read more about L2D on the official Huggingface blog: LeRobot goes to driving school 90+ TeraBytes of multimodal data (5000+ hours of driving) from 30 cities in Germany 6x surrounding HD cameras and complete vehicle state: Speed/Heading/GPS/IMU Continuous: Gas/Brake/Steering and discrete actions: Gear/Turn Signals Environment state: Lane count, Road type (highway|residential), Road surface (asphalt, cobbled, sett), Max speed limit.… See the full description on the dataset page: https://huggingface.co/datasets/yaak-ai/L2D.tabularrobotics10M<n<100M52 likes72k downloads4mo agoHugging Face15Fsoft-AIC /RobotDesign1M RobotDesign1M: A Large-scale Dataset for Robot Design Understanding RobotDesign1M is a large-scale, multimodal dataset for robot design understanding, built from image–text data curated from scientific literature across a wide range of robotics domains. It is designed to support research on design-aware foundation models, including design image generation, visual question answering about designs, and design image retrieval. 📄 Paper: RobotDesign1M: A Large-scale Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/Fsoft-AIC/RobotDesign1M.imageimage-text-to-text1M<n<10M8 likes70k downloads2mo agoHugging Face16markov-ai /cad-environments CAD Environments CAD Environments is a multimodal dataset of complete, human-performed workflows in desktop CAD software. The current release contains 51 task workflows totaling 99.03 hours, covering eight software groups across mechanical design, architecture, MEP, structural design, and general 3D modeling. Each workflow preserves the full task context—not just the final model—including the problem statement, reference and input files, a gold output, evaluation rubrics, a… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/cad-environments.imagen<1K17 likes57k downloads2mo agoHugging Face17applied-ai-018 /pretraining_v1-omega5 likes56k downloads2y agoHugging Face18fixie-ai /covost2This is a partial copy of CoVoST2 dataset. The main difference is that the audio data is included in the dataset, which makes usage easier and allows browsing the samples using HF Dataset Viewer. The limitation of this method is that all audio samples of the EN_XX subsets are duplicated, as such the size of the dataset is larger. As such, not all the data is included: Only the validation and test subsets are available. From the XX_EN subsets, only fr, es, and zh-CN are included. audio1M<n<10M5 likes55k downloads2y agoHugging Face19HuggingFaceH4 /aime_2024 Dataset card for AIME 2024 This dataset consists of 30 problems from the 2024 AIME I and AIME II tests. The original source is AI-MO/aimo-validation-aime, which contains a larger set of 90 problems from AIME 2022-2024. textn<1K64 likes55k downloads2y agoHugging Face20ai-for-good-lab /ai4g-flood-dataset Flood Detection Dataset Introduction This dataset accompanies the paper Mapping global floods with 10 years of satellite radar data (Nature Communications, 2025) and contains global flood detections derived from Sentinel-1 Synthetic Aperture Radar (SAR) imagery using a deep learning change detection model. The dataset spans October 2014 – September 2024, offering a longitudinal view of flood-prone areas worldwide. Key features: Cloud-penetrating SAR data for consistent… See the full description on the dataset page: https://huggingface.co/datasets/ai-for-good-lab/ai4g-flood-dataset.imagen<1K16 likes52k downloads11mo agoHugging Face21AI-MO /NuminaMath-1.5 Dataset Card for NuminaMath 1.5 Dataset Summary This is the second iteration of the popular NuminaMath dataset, bringing high quality post-training data for approximately 900k competition-level math problems. Each solution is formatted in a Chain of Thought (CoT) manner. The sources of the dataset range from Chinese high school math exercises to US and international mathematics olympiad competition problems. The data were primarily collected from online exam paper PDFs… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/NuminaMath-1.5.texttext-generation100K<n<1M194 likes49k downloads8mo agoHugging Face22skylenage-ai /HLE-Verified HLE-Verified A Systematic Verification and Structured Revision of Humanity’s Last Exam Overview Humanity’s Last Exam (HLE) is a high-difficulty, multi-domain benchmark designed to evaluate advanced reasoning capabilities across diverse scientific and technical domains. Following its public release, members of the open-source community raised concerns regarding the reliability of certain items. Community discussions and informal replication attempts suggested that some… See the full description on the dataset page: https://huggingface.co/datasets/skylenage-ai/HLE-Verified.document1K<n<10K19 likes45k downloads7mo agoHugging Face23markov-ai /computer-use-large Computer Use Large A large-scale dataset of 48,478 screen recording videos (~12,300 hours) of professional software being used, sourced from the internet. All videos have been trimmed to remove non-screen-recording content (intros, outros, talking heads, transitions) and audio has been stripped. Dataset Summary Category Videos Hours AutoCAD 10,059 2,149 Blender 11,493 3,624 Excel 8,111 2,002 Photoshop 10,704 2,060 Salesforce 7,807 2,336 VS… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/computer-use-large.tabularvideo-classification10K<n<100K196 likes45k downloads6mo agoHugging Face24airtrain-ai /fineweb-edu-fortified Fineweb-Edu-Fortified The composition of fineweb-edu-fortified, produced by automatically clustering a 500k row sample in Airtrain What is it? Fineweb-Edu-Fortified is a dataset derived from Fineweb-Edu by applying exact-match deduplication across the whole dataset and producing an embedding for each row. The number of times the text from each row appears is also included as a count column. The embeddings were produced using TaylorAI/bge-micro Fineweb and… See the full description on the dataset page: https://huggingface.co/datasets/airtrain-ai/fineweb-edu-fortified.tabulartext-generation100M<n<1B65 likes43k downloads2y agoHugging Face25ai4bharat /sangraha Sangraha Sangraha is the largest high-quality, cleaned Indic language pretraining data containing 251B tokens summed up over 22 languages, extracted from curated sources, existing multilingual corpora and large scale translations. More information: For detailed information on the curation and cleaning process of Sangraha, please checkout our paper on Arxiv; Check out the scraping and cleaning pipelines used to curate Sangraha on GitHub; Getting Started For… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/sangraha.texttext-generation100M<n<1B83 likes42k downloads2y agoHugging Face26lmarena-ai /leaderboard-dataset Arena Leaderboard Dataset Historical snapshots of the Arena leaderboard. Usage from datasets import load_dataset # Load all historical text style control data ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="full") # Load the current text style control leaderboard ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="latest") # Filter to overall category ds =… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset.tabular1M<n<10M25 likes41k downloads6d agoHugging Face27AI-MO /aimo-validation-aime Dataset Card for AIMO Validation AIME All 90 problems come from AIME 22, AIME 23, and AIME 24, and have been extracted directly from the AOPS wiki page https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions This dataset serves as an internal validation set during our participation in the AIMO progress prize competition. Using data after 2021 is to avoid potential overlap with the MATH training set. Here are the different columns in the dataset: problem: the… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/aimo-validation-aime.textn<1K68 likes37k downloads1y agoHugging Face28yentinglin /aime_2025 AIME 2025 This dataset contains 30 problems from the 2025 AIME tests, including: AIME I: 15 problems AIME II: 15 problems tabularn<1K12 likes37k downloads9mo agoHugging Face29Robeedau /airlens-live AirLens Live Data Live data layer for AirLens, an open air-quality monitoring platform. Updated by scheduled GitHub Actions pipelines. Layout mirrors the former Supabase Storage buckets: Path Content Cadence aq-data/current-*-grid.json Global pollutant grids (PM2.5/PM10/O3/NO2/CO) hourly aq-data/timeline/ GEFS-Aerosols PM2.5 frames, -24h..+24h, 3h step every 3h aq-data/predictions/grid_latest.json AOD→PM2.5 model predictions (p10-p90 + DQSS) every 3h… See the full description on the dataset page: https://huggingface.co/datasets/Robeedau/airlens-live.2 likes36k downloads23m agoHugging Face30MathArena /aime_2026 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from AIME 2026 used for the MathArena Leaderboard Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. answer (int64): Gold final answer. problem (string): Problem statement, usually stored as LaTeX source. Source… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/aime_2026.tabularn<1K61 likes35k downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.