CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /ai2_arc Dataset Card for "ai2_arc" Dataset Summary A new dataset of 7,787 genuine grade-school level, multiple-choice science questions, assembled to encourage research in advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. We are also including a corpus of over 14 million science sentences… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ai2_arc.textquestion-answering1K<n<10K402 likes875k downloads3y agoHugging Face02SakanaAI /AI-CUDA-Engineer-Archive The AI CUDA Engineer Archive 👷: Agentic CUDA Kernel Discovery, Optimization & Composition We release The AI CUDA Engineer archive, a dataset consisting of approximately 30,000 CUDA kernels generated by The AI CUDA Engineer. It is released under the CC-By-4.0 license and can be accessed via HuggingFace and interactively visualized here. The dataset is based on the Kernel tasks provided in KernelBench and includes a torch reference implementation, torch, NCU and Clang-tidy… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/AI-CUDA-Engineer-Archive.tabular10K<n<100K227 likes136k downloads2y agoHugging Face03picbreeder-vlm /picbreeder-vlm-archive Picbreeder-VLM Archive Every image evolved by the swarm of vision-language-model "breeders" in In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models (GECCO 2026), together with the CPPN genomes that produced them, the agents' reasoning transcripts, the lineage graphs, and the analysis artifacts behind the paper and blog. The original Picbreeder (Secretan et al., 2008) let crowds of humans collaboratively evolve images from CPPN… See the full description on the dataset page: https://huggingface.co/datasets/picbreeder-vlm/picbreeder-vlm-archive.imageimage-to-text100K<n<1M14 likes73k downloads2mo agoHugging Face04arcinstitute /State-Parse-FilteredThe single cell RNA-seq dataset with human PBMC samples was sourced from Parse Biosciences [1]. [1] Performance of Evercode™ WT v3 in Human Immune Cells (PBMCs), https://www.parsebiosciences.com/datasets/performance-of-evercode-wt-v3-in-human-immune-cells-pbmcs/; Parse Biosciences, Seattle, USA; accessed 05/27/2025. Certain uses of this data may require a license from Parse Biosciences, Inc. textn<1K0 likes65k downloads4mo agoHugging Face05farhanhubble /jfk-archives Dataset Card for JFK Archives This dataset is a collection of all records pertaining to the assassination of the US president, John F. Kennedy, released until April 2025 through archives.org by the US government. Dataset Details Dataset Description The original data downloaded from archives.org consists of 56,300 scanned documents in PDF format, released until April 2025. The files are organized by their release year(s): 2107-2018, 2021, 2022, 2023 and 2025.… See the full description on the dataset page: https://huggingface.co/datasets/farhanhubble/jfk-archives.textquestion-answering10K<n<100K0 likes45k downloads1y agoHugging Face06arcee-ai /distilabel-intel-orca-dpo-pairs-binarizedThis is the binarized version of distilabel Orca Pairs for DPO and ORPO. Reference: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs?row=0 text10K<n<100K1 likes24k downloads2y agoHugging Face07aicrowd /arc-whestbench-public-2026 Organized by: Alignment Research Center (ARC), AIcrowd WhestBench 2026: ARC White-Box Estimation Challenge WhestBench is a benchmark for white-box activation estimation: given the weights of a randomly initialized ReLU multi-layer perceptron (MLP) and a strict floating-point-operation (FLOP) budget, predict the average post-activation value of every neuron when the network is fed standard Gaussian inputs. This is the WhestBench 2026… See the full description on the dataset page: https://huggingface.co/datasets/aicrowd/arc-whestbench-public-2026.tabularother1K<n<10K0 likes21k downloads24d agoHugging Face08Matthijs /cmu-arctic-xvectors Speaker embeddings extracted from CMU ARCTIC There is one .npy file for each utterance in the dataset, 7931 files in total. The speaker embeddings are 512-element X-vectors. The CMU ARCTIC dataset divides the utterances among the following speakers: bdl (US male) slt (US female) jmk (Canadian male) awb (Scottish male) rms (US male) clb (US female) ksp (Indian male) The X-vectors were extracted using this script, which uses the speechbrain/spkrec-xvect-voxceleb model. Usage: from… See the full description on the dataset page: https://huggingface.co/datasets/Matthijs/cmu-arctic-xvectors.texttext-to-speech1K<n<10K64 likes19k downloads4y agoHugging Face09ArchEGraph /ArchEGraph ArchEGraph ArchEGraph is a building-energy dataset organized for graph-based and weather-conditioned learning. Dataset Summary Total cases in manifest.csv: 49,326 Unique buildings: 5,481 Unique weather IDs: 64 n_steps range: 968 to 8,760 n_spaces range: 1 to 231 This repository currently stores: manifest.csv (index of all cases) building/ (5,481 files) geometry/ (5,482 files) weather/ (64 files) energy/ (49,326 files; nested under subfolders like 00/) split/… See the full description on the dataset page: https://huggingface.co/datasets/ArchEGraph/ArchEGraph.tabulargraph-ml100K<n<1M1 likes9.6k downloads5mo agoHugging Face10alexandrainst /m_arc Multilingual ARC Dataset Summary This dataset is a machine translated version of the ARC dataset. The Icelandic (is) part was translated with Miðeind's Greynir model and Norwegian (nb) was translated with DeepL. The rest of the languages was translated using GPT-3.5-turbo by the University of Oregon, and this part of the dataset was originally uploaded to this Github repository. textquestion-answering10K<n<100K4 likes9.2k downloads3y agoHugging Face11AiAF /SCPWiki-Cleaned-PDF-Archivesdocumenttext-generationn<1K1 likes7.7k downloads1y agoHugging Face12ust-archive /scheduleSee https://github.com/ust-archive/ust-archive for more information. tabular100K<n<1M0 likes7.5k downloads4h agoHugging Face13gavinlaw /rl-run-archive-2026 RL run archive 2026 Archived raw run artifacts (rollout trajectories, rendered frames, policy and optimizer checkpoints, configs, logs) from simulation reinforcement-learning experiments, published for long-term preservation and reproducibility. Layout mirrors the verified backup trees they were copied from: tilde/20260915-102000/ and taurus/20260915-085631/: batched tar archives. Every archive carries a per-file SHA-256 manifest inside it; the batch inventories (9998.json.gz… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/rl-run-archive-2026.tabularn<1K0 likes6.6k downloads2d agoHugging Face14AiAF /SCPWiki-Archive-02-March-2025-Datasetstextn<1K0 likes6.6k downloads2y agoHugging Face15arcprize /arc_agi_v2_public_evaltext10K<n<100K7 likes5.2k downloads4mo agoHugging Face16CoreEmotionFramework /CEF_Main_Archive Core Emotion Framework (CEF) Main Archive The Decalogue of Operators The Core Emotion Framework defines exactly ten functional operators. This is the complete and authoritative set. No additional operators exist. No operators may be removed, renamed, or substituted. This dataset serves as the absolute source of truth for the following: Sensing Calculating Deciding Expanding Constricting Achieving Arranging Appreciating Boosting Accepting { "@context":… See the full description on the dataset page: https://huggingface.co/datasets/CoreEmotionFramework/CEF_Main_Archive.documentothern<1K0 likes4.7k downloads3mo agoHugging Face17Memento-ARC /ultimatedocument0 likes4.3k downloads2mo agoHugging Face18common-pile /github_archive GitHub Archive Description According to GitHub’s terms of service, issues and pull request descriptions—along with the their comments—inherit the license of their associated repository. To collect this data, we used the GitHub Archive’s public BigQuery table of events to extracted all issue, pull request, and comment events since 2011 and aggregated them into threads. The table appeared to be missing “edit” events so the text from each comment is the original from when… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive.texttext-generation10M<n<100M2 likes4.2k downloads1y agoHugging Face19benjamin-paine /free-music-archive-full FMA: A Dataset for Music Analysis Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson. International Society for Music Information Retrieval Conference (ISMIR), 2017. We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-full.audioaudio-to-audio100K<n<1M20 likes3.5k downloads2y agoHugging Face20hynky /okapi_arc_challengetext10K<n<100K0 likes2.7k downloads2y agoHugging Face21eastbrush /eastbrush_archive Eastbrush Archive Official Website (Full Archive System): https://www.eastbrush.com This dataset contains high-resolution images and structured tags for AI training. The full archive system — including chapter exhibitions, structural context, and extended records — is available on the official website. What is my true self? The Eastbrush Archive is a long-term, evolving system that documents the visual language ofJang Byeong Eun (Eastbrush / 張炳彥) — a painter whose… See the full description on the dataset page: https://huggingface.co/datasets/eastbrush/eastbrush_archive.imagen<1K2 likes2.5k downloads3d agoHugging Face22Snowflake /msmarco-v2.1-snowflake-arctic-embed-l Snowflake Arctic Embed L Embeddings for MSMARCO V2.1 for TREC-RAG This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG All embeddings are created using Snowflake's Arctic Embed L and are intended to serve as a simple baseline for dense retrieval-based methods. Retrieval Performance Retrieval performance for the TREC DL21-23, MSMARCOV2-Dev and Raggy Queries can be found below with BM25 as a baseline. For both… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-l.textquestion-answering10M<n<100M0 likes2.5k downloads2y agoHugging Face23davanstrien /prelinger-archives-open Prelinger Archives Open License Videos A collection of historical films from the Prelinger Archives on the Internet Archive, filtered to include only videos with open licenses (Public Domain, CC0, CC BY, CC BY-SA). Dataset Description The Prelinger Archives is a collection of over 17,000 advertising, educational, industrial, and amateur films. This dataset contains the subset of videos that are available under open licenses, making them freely usable for research… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/prelinger-archives-open.textvideo-classification1K<n<10K1 likes2.3k downloads6mo agoHugging Face24leesharks /crimson-hexagonal-archive The Crimson Hexagonal Archive — machine-readable representation Query this without downloading anything. Every config is served by the Hugging Face datasets-server over plain HTTP, no auth, no client library. Use /rows — it is the reliable one. It reads the parquet directly and answers in under two seconds: https://datasets-server.huggingface.co/rows?dataset=leesharks%2Fcrimson-hexagonal-archive&config=deposits&split=train&offset=0&length=10… See the full description on the dataset page: https://huggingface.co/datasets/leesharks/crimson-hexagonal-archive.tabular10K<n<100K2 likes2.3k downloads3h agoHugging Face25ThomasTheMaker /Arc-Corpustext10M<n<100M0 likes2.1k downloads10mo agoHugging Face26simon123905 /trellis500k-sketchfab-archivestabularn<1K0 likes2k downloads6mo agoHugging Face27arcee-globe /ACVA-10percenttext1K<n<10K0 likes2k downloads2y agoHugging Face28nvidia /Nemotron-SFT-ARC-AGI-v1 Dataset Description: Nemotron-SFT-ARC-AGI-v1 is a supervised fine-tuning (SFT) dataset of multi-turn agentic reasoning traces produced by open-weight large language models attempting to solve ARC-AGI visual-reasoning puzzles. Each ARC puzzle (a set of (input grid, output grid) demonstration pairs plus one or more test inputs, where grids are 2D integer arrays representing colors) is formatted as a text prompt and given to an agent powered by one of nine open-weight reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-ARC-AGI-v1.texttext-generation100K<n<1M22 likes1.9k downloads4mo agoHugging Face29SimulaMet /moltbook-observatory-archive Observatory Dataset This dataset is an incremental export of a SQLite observatory database, published as date-partitioned Parquet files for efficient browsing and querying on Hugging Face. For example, you can filter data by wildcards on date: ds = load_dataset( "SimulaMet/moltbook-observatory-archive", "posts", data_files="data/posts/2026-01-2*.parquet", # 20–29 split="train" ) Each SQLite table is exposed as a separate dataset subset. Use dropdown above the… See the full description on the dataset page: https://huggingface.co/datasets/SimulaMet/moltbook-observatory-archive.tabular1M<n<10M32 likes1.9k downloads11d agoHugging Face30alexdum /meteogate-archive Meteogate European Weather Observations Archive A continuously growing archive of real-time meteorological observations from the EUMETNET Meteogate E-SOH service, covering thousands of weather stations across Europe. Data Structure Each Parquet file contains observations in long format (one row per station × variable × timestamp) with the following columns: Column Type Description timestamp datetime Observation time in UTC station_id string WIGOS… See the full description on the dataset page: https://huggingface.co/datasets/alexdum/meteogate-archive.tabulartime-series-forecasting100M<n<1B0 likes1.9k downloads10h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.