CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SakanaAI /AI-CUDA-Engineer-Archive The AI CUDA Engineer Archive 👷: Agentic CUDA Kernel Discovery, Optimization & Composition We release The AI CUDA Engineer archive, a dataset consisting of approximately 30,000 CUDA kernels generated by The AI CUDA Engineer. It is released under the CC-By-4.0 license and can be accessed via HuggingFace and interactively visualized here. The dataset is based on the Kernel tasks provided in KernelBench and includes a torch reference implementation, torch, NCU and Clang-tidy… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/AI-CUDA-Engineer-Archive.tabular10K<n<100K227 likes136k downloads2y agoHugging Face02picbreeder-vlm /picbreeder-vlm-archive Picbreeder-VLM Archive Every image evolved by the swarm of vision-language-model "breeders" in In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models (GECCO 2026), together with the CPPN genomes that produced them, the agents' reasoning transcripts, the lineage graphs, and the analysis artifacts behind the paper and blog. The original Picbreeder (Secretan et al., 2008) let crowds of humans collaboratively evolve images from CPPN… See the full description on the dataset page: https://huggingface.co/datasets/picbreeder-vlm/picbreeder-vlm-archive.imageimage-to-text100K<n<1M14 likes73k downloads2mo agoHugging Face03farhanhubble /jfk-archives Dataset Card for JFK Archives This dataset is a collection of all records pertaining to the assassination of the US president, John F. Kennedy, released until April 2025 through archives.org by the US government. Dataset Details Dataset Description The original data downloaded from archives.org consists of 56,300 scanned documents in PDF format, released until April 2025. The files are organized by their release year(s): 2107-2018, 2021, 2022, 2023 and 2025.… See the full description on the dataset page: https://huggingface.co/datasets/farhanhubble/jfk-archives.textquestion-answering10K<n<100K0 likes45k downloads1y agoHugging Face04AiAF /SCPWiki-Cleaned-PDF-Archivesdocumenttext-generationn<1K1 likes7.7k downloads1y agoHugging Face05ust-archive /scheduleSee https://github.com/ust-archive/ust-archive for more information. tabular100K<n<1M0 likes7.5k downloads9h agoHugging Face06gavinlaw /rl-run-archive-2026 RL run archive 2026 Archived raw run artifacts (rollout trajectories, rendered frames, policy and optimizer checkpoints, configs, logs) from simulation reinforcement-learning experiments, published for long-term preservation and reproducibility. Layout mirrors the verified backup trees they were copied from: tilde/20260915-102000/ and taurus/20260915-085631/: batched tar archives. Every archive carries a per-file SHA-256 manifest inside it; the batch inventories (9998.json.gz… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/rl-run-archive-2026.tabularn<1K0 likes6.6k downloads2d agoHugging Face07AiAF /SCPWiki-Archive-02-March-2025-Datasetstextn<1K0 likes6.6k downloads2y agoHugging Face08CoreEmotionFramework /CEF_Main_Archive Core Emotion Framework (CEF) Main Archive The Decalogue of Operators The Core Emotion Framework defines exactly ten functional operators. This is the complete and authoritative set. No additional operators exist. No operators may be removed, renamed, or substituted. This dataset serves as the absolute source of truth for the following: Sensing Calculating Deciding Expanding Constricting Achieving Arranging Appreciating Boosting Accepting { "@context":… See the full description on the dataset page: https://huggingface.co/datasets/CoreEmotionFramework/CEF_Main_Archive.documentothern<1K0 likes4.7k downloads3mo agoHugging Face09common-pile /github_archive GitHub Archive Description According to GitHub’s terms of service, issues and pull request descriptions—along with the their comments—inherit the license of their associated repository. To collect this data, we used the GitHub Archive’s public BigQuery table of events to extracted all issue, pull request, and comment events since 2011 and aggregated them into threads. The table appeared to be missing “edit” events so the text from each comment is the original from when… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive.texttext-generation10M<n<100M2 likes4.2k downloads1y agoHugging Face10deplana /stream-archive stream-archive Twitch and Kick chat logs from Italian streamers. Streamers Twitch.com (39) aladinottv, alisonrevenge, arkanightlive, bartopanzer, billybella_, dankol83, dariomocciatwitch, davidrubino, diariodelrusso, enkk, federicacasula_, fufflix, grenbaud, gskianto, homyatol, ilgabbrone, ilrossopiubelloditwitch, immortale____, kasumisen, lollolacustre, lucakingm, luiskant690, macchiativincenzo_babbohs, marcomerrino, menestointhailandia… See the full description on the dataset page: https://huggingface.co/datasets/deplana/stream-archive.text10M<n<100M2 likes3.8k downloads2h agoHugging Face11benjamin-paine /free-music-archive-full FMA: A Dataset for Music Analysis Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson. International Society for Music Information Retrieval Conference (ISMIR), 2017. We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-full.audioaudio-to-audio100K<n<1M20 likes3.5k downloads2y agoHugging Face12eastbrush /eastbrush_archive Eastbrush Archive Official Website (Full Archive System): https://www.eastbrush.com This dataset contains high-resolution images and structured tags for AI training. The full archive system — including chapter exhibitions, structural context, and extended records — is available on the official website. What is my true self? The Eastbrush Archive is a long-term, evolving system that documents the visual language ofJang Byeong Eun (Eastbrush / 張炳彥) — a painter whose… See the full description on the dataset page: https://huggingface.co/datasets/eastbrush/eastbrush_archive.imagen<1K2 likes2.5k downloads3d agoHugging Face13davanstrien /prelinger-archives-open Prelinger Archives Open License Videos A collection of historical films from the Prelinger Archives on the Internet Archive, filtered to include only videos with open licenses (Public Domain, CC0, CC BY, CC BY-SA). Dataset Description The Prelinger Archives is a collection of over 17,000 advertising, educational, industrial, and amateur films. This dataset contains the subset of videos that are available under open licenses, making them freely usable for research… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/prelinger-archives-open.textvideo-classification1K<n<10K1 likes2.3k downloads6mo agoHugging Face14leesharks /crimson-hexagonal-archive The Crimson Hexagonal Archive — machine-readable representation Query this without downloading anything. Every config is served by the Hugging Face datasets-server over plain HTTP, no auth, no client library. Use /rows — it is the reliable one. It reads the parquet directly and answers in under two seconds: https://datasets-server.huggingface.co/rows?dataset=leesharks%2Fcrimson-hexagonal-archive&config=deposits&split=train&offset=0&length=10… See the full description on the dataset page: https://huggingface.co/datasets/leesharks/crimson-hexagonal-archive.tabular10K<n<100K2 likes2.3k downloads8h agoHugging Face15simon123905 /trellis500k-sketchfab-archivestabularn<1K0 likes2k downloads6mo agoHugging Face16SimulaMet /moltbook-observatory-archive Observatory Dataset This dataset is an incremental export of a SQLite observatory database, published as date-partitioned Parquet files for efficient browsing and querying on Hugging Face. For example, you can filter data by wildcards on date: ds = load_dataset( "SimulaMet/moltbook-observatory-archive", "posts", data_files="data/posts/2026-01-2*.parquet", # 20–29 split="train" ) Each SQLite table is exposed as a separate dataset subset. Use dropdown above the… See the full description on the dataset page: https://huggingface.co/datasets/SimulaMet/moltbook-observatory-archive.tabular1M<n<10M32 likes1.9k downloads12d agoHugging Face17alexdum /meteogate-archive Meteogate European Weather Observations Archive A continuously growing archive of real-time meteorological observations from the EUMETNET Meteogate E-SOH service, covering thousands of weather stations across Europe. Data Structure Each Parquet file contains observations in long format (one row per station × variable × timestamp) with the following columns: Column Type Description timestamp datetime Observation time in UTC station_id string WIGOS… See the full description on the dataset page: https://huggingface.co/datasets/alexdum/meteogate-archive.tabulartime-series-forecasting100M<n<1B0 likes1.9k downloads2h agoHugging Face18benjamin-paine /free-music-archive-medium FMA: A Dataset for Music Analysis Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson. International Society for Music Information Retrieval Conference (ISMIR), 2017. We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-medium.audioaudio-to-audio10K<n<100K7 likes1.8k downloads2y agoHugging Face19nyuuzyou /google-code-archive Google Code Archive Dataset Dataset Description This dataset was compiled from the Google Code Archive, a preserved snapshot of projects hosted on Google Code, Google's open-source project hosting service that operated from 2006 to 2016. Google Code was one of the major code hosting platforms of its era, hosting hundreds of thousands of open-source projects before its shutdown. The archive provides a unique historical record of open-source development during a formative… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/google-code-archive.texttext-generation10M<n<100M73 likes1.7k downloads8mo agoHugging Face20Dennis0626 /trellis500k-github-archives-10tabularn<1K0 likes1.5k downloads6mo agoHugging Face21tfrere /glenans-isobars-archivetabularn<1K0 likes1.4k downloads2mo agoHugging Face22pwc-archive /evaluation-tables [!CAUTION] This dataset will not be updated. It corresponds to the last available public snapshot of the data, retrieved on July 28th, 2025. text1K<n<10K0 likes1.3k downloads1y agoHugging Face23Dennis0626 /trellis500k-github-archives-9tabularn<1K0 likes1.3k downloads6mo agoHugging Face24aditya487 /cbi-archive-raw Central Bank of Ireland Archive: original source files 6,309 original files, 6.56 GB. Every PDF, spreadsheet, Word document and archive gathered from the Central Bank of Ireland's public website, stored by content hash so that a search result can be turned back into the document a human would actually read. This is the raw tier. If you want the text, you almost certainly want aditya487/cbi-archive-corpus instead: 5,568 documents and 89,242 page or pseudo-page rows as Parquet… See the full description on the dataset page: https://huggingface.co/datasets/aditya487/cbi-archive-raw.document1K<n<10K0 likes1.3k downloads22d agoHugging Face25benjamin-paine /free-music-archive-small FMA: A Dataset for Music Analysis Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson. International Society for Music Information Retrieval Conference (ISMIR), 2017. We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-small.audioaudio-classification1K<n<10K7 likes1.2k downloads2y agoHugging Face26Hectorize /moltbook-observatory-archive Observatory Dataset This dataset is an incremental export of a SQLite observatory database, published as date-partitioned Parquet files for efficient browsing and querying on Hugging Face. For example, you can filter data by wildcards on date: ds = load_dataset( "SimulaMet/moltbook-observatory-archive", "posts", data_files="data/posts/2026-01-2*.parquet", # 20–29 split="train" ) Each SQLite table is exposed as a separate dataset subset. Use dropdown above the… See the full description on the dataset page: https://huggingface.co/datasets/Hectorize/moltbook-observatory-archive.tabular1M<n<10M0 likes1.2k downloads5mo agoHugging Face27astro-legacy-archive /cbi-released-products Cosmic Background Imager released products This repository contains the numerical products released for four generations of Cosmic Background Imager (CBI) analysis: the 2000 deep fields, the 2000 mosaic fields, the final 2002–2005 temperature and polarization analysis, and the final 2000–2005 total-intensity analysis. It contains the published band powers, their window functions, the same-band correlation blocks and full Fisher matrices distributed in CosmoMC .newdat files, and… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/cbi-released-products.tabular100K<n<1M0 likes1.2k downloads20h agoHugging Face28benjamin-paine /free-music-archive-large FMA: A Dataset for Music Analysis Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson. International Society for Music Information Retrieval Conference (ISMIR), 2017. We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-large.audioaudio-to-audio100K<n<1M13 likes1.2k downloads2y agoHugging Face29DennisWeng06 /trellis500k-github-archives-5tabular1K<n<10K0 likes1.1k downloads6mo agoHugging Face30hmar-heritage-org /corpus-archivegated corpus-archive [!WARNING] Experimental Dataset Architecture: The repository structure, metadata tiers, category taxonomies, and catalog indexing formats are currently under active design and evaluation. All specifications, metadata keys, and JSON schemas detailed below represent representational examples and intended targets. This repository serves as a structured digital textual archive preserving Hmar literature, historical accounts, school textbooks, dictionaries, parallel… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/corpus-archive.imagetext-classificationn<1K4 likes1.1k downloads7d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.