CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01banned-historical-archives /banned-historical-archives 和谐历史档案馆数据集 - Banned Historical Archives Datasets 和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。 目录结构 banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步 raw # 原始文件 config # 配置文件 todo # 存放暂未录入网站的文件 部分报纸和图片资料存放在单独的仓库: 名称 地址 状态 参考消息 https://huggingface.co/datasets/banned-historical-archives/ckxx 未录入 人民日报 https://huggingface.co/datasets/banned-historical-archives/rmrb 已精选重要的文章录入 文汇报… See the full description on the dataset page: https://huggingface.co/datasets/banned-historical-archives/banned-historical-archives.imagen<1K91 likes1.7m downloads11mo agoHugging Face02allenai /ai2_arc Dataset Card for "ai2_arc" Dataset Summary A new dataset of 7,787 genuine grade-school level, multiple-choice science questions, assembled to encourage research in advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. We are also including a corpus of over 14 million science sentences… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ai2_arc.textquestion-answering1K<n<10K402 likes871k downloads3y agoHugging Face03banned-historical-archives /zhongyangribao2 likes220k downloads2y agoHugging Face04Metanova /Submission-Archive1 likes192k downloads24m agoHugging Face05SakanaAI /AI-CUDA-Engineer-Archive The AI CUDA Engineer Archive 👷: Agentic CUDA Kernel Discovery, Optimization & Composition We release The AI CUDA Engineer archive, a dataset consisting of approximately 30,000 CUDA kernels generated by The AI CUDA Engineer. It is released under the CC-By-4.0 license and can be accessed via HuggingFace and interactively visualized here. The dataset is based on the Kernel tasks provided in KernelBench and includes a torch reference implementation, torch, NCU and Clang-tidy… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/AI-CUDA-Engineer-Archive.tabular10K<n<100K227 likes136k downloads2y agoHugging Face06open-index /arctic Arctic Shift Reddit Archive Every Reddit comment and submission since 2005, organized as monthly Parquet shards What is it? The full Reddit archive from Arctic Shift, converted to Parquet and hosted here for easy access. Covers every public subreddit from 2005-12 through 2026-02. Right now the archive has 15.7B items (12.9B comments, 2.8B submissions) in 1.3 TB of compressed Parquet. Comments and submissions are stored as separate datasets, split into monthly… See the full description on the dataset page: https://huggingface.co/datasets/open-index/arctic.text-generation1B<n<10B29 likes98k downloads2mo agoHugging Face07picbreeder-vlm /picbreeder-vlm-archive Picbreeder-VLM Archive Every image evolved by the swarm of vision-language-model "breeders" in In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models (GECCO 2026), together with the CPPN genomes that produced them, the agents' reasoning transcripts, the lineage graphs, and the analysis artifacts behind the paper and blog. The original Picbreeder (Secretan et al., 2008) let crowds of humans collaboratively evolve images from CPPN… See the full description on the dataset page: https://huggingface.co/datasets/picbreeder-vlm/picbreeder-vlm-archive.imageimage-to-text100K<n<1M14 likes74k downloads2mo agoHugging Face08arcinstitute /State-Parse-FilteredThe single cell RNA-seq dataset with human PBMC samples was sourced from Parse Biosciences [1]. [1] Performance of Evercode™ WT v3 in Human Immune Cells (PBMCs), https://www.parsebiosciences.com/datasets/performance-of-evercode-wt-v3-in-human-immune-cells-pbmcs/; Parse Biosciences, Seattle, USA; accessed 05/27/2025. Certain uses of this data may require a license from Parse Biosciences, Inc. textn<1K0 likes65k downloads4mo agoHugging Face09farhanhubble /jfk-archives Dataset Card for JFK Archives This dataset is a collection of all records pertaining to the assassination of the US president, John F. Kennedy, released until April 2025 through archives.org by the US government. Dataset Details Dataset Description The original data downloaded from archives.org consists of 56,300 scanned documents in PDF format, released until April 2025. The files are organized by their release year(s): 2107-2018, 2021, 2022, 2023 and 2025.… See the full description on the dataset page: https://huggingface.co/datasets/farhanhubble/jfk-archives.textquestion-answering10K<n<100K0 likes46k downloads1y agoHugging Face10arcinstitute /opengenome2 OpenGenome2 OpenGenome2 is a database of nearly 9 trillion base pairs of curated DNA from across all domains of life. Collected from diverse species and public data sources, OpenGenome2 was used to train Evo 2 models. Please refer to the Evo 2 preprint or github repository for further details and usage examples. We provide OpenGenome2 in two formats, the dataset is organized into two main directories to reflect this: fasta which contain the DNA sequences jsonl which… See the full description on the dataset page: https://huggingface.co/datasets/arcinstitute/opengenome2.text-generationn>1T159 likes39k downloads19d agoHugging Face11banned-historical-archives /low-priority1 likes33k downloads2y agoHugging Face12arcee-ai /distilabel-intel-orca-dpo-pairs-binarizedThis is the binarized version of distilabel Orca Pairs for DPO and ORPO. Reference: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs?row=0 text10K<n<100K1 likes24k downloads2y agoHugging Face13aicrowd /arc-whestbench-public-2026 Organized by: Alignment Research Center (ARC), AIcrowd WhestBench 2026: ARC White-Box Estimation Challenge WhestBench is a benchmark for white-box activation estimation: given the weights of a randomly initialized ReLU multi-layer perceptron (MLP) and a strict floating-point-operation (FLOP) budget, predict the average post-activation value of every neuron when the network is fed standard Gaussian inputs. This is the WhestBench 2026… See the full description on the dataset page: https://huggingface.co/datasets/aicrowd/arc-whestbench-public-2026.tabularother1K<n<10K0 likes21k downloads23d agoHugging Face14Matthijs /cmu-arctic-xvectors Speaker embeddings extracted from CMU ARCTIC There is one .npy file for each utterance in the dataset, 7931 files in total. The speaker embeddings are 512-element X-vectors. The CMU ARCTIC dataset divides the utterances among the following speakers: bdl (US male) slt (US female) jmk (Canadian male) awb (Scottish male) rms (US male) clb (US female) ksp (Indian male) The X-vectors were extracted using this script, which uses the speechbrain/spkrec-xvect-voxceleb model. Usage: from… See the full description on the dataset page: https://huggingface.co/datasets/Matthijs/cmu-arctic-xvectors.texttext-to-speech1K<n<10K64 likes19k downloads4y agoHugging Face15arcinstitute /Stack-scBaseCount189M0 likes17k downloads4mo agoHugging Face16Exploration-Lab /dim-discovery-archive Geometry of Decision Making in Language Models Abhinav Joshi · Divyanshu Bhatt · Ashutosh ModiNeurIPS 2025 This repository contains the official implementation/release for the NeurIPS 2025 paper Geometry of Decision Making in Language Models. We study the internal decision-making processes of large language models through the lens of intrinsic dimension (ID), analyzing how hidden representations evolve across layers in a multiple-choice… See the full description on the dataset page: https://huggingface.co/datasets/Exploration-Lab/dim-discovery-archive.0 likes16k downloads8mo agoHugging Face17Dragonegg2026 /banned-historical-archives 和谐历史档案馆数据集 - Banned Historical Archives Datasets 和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。 目录结构 banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步 raw # 原始文件 config # 配置文件 todo # 存放暂未录入网站的文件 部分报纸和图片资料存放在单独的仓库: 名称 地址 状态 参考消息 https://huggingface.co/datasets/banned-historical-archives/ckxx 未录入 人民日报 https://huggingface.co/datasets/banned-historical-archives/rmrb 已精选重要的文章录入 文汇报… See the full description on the dataset page: https://huggingface.co/datasets/Dragonegg2026/banned-historical-archives.imagen<1K1 likes15k downloads5mo agoHugging Face18banned-historical-archives /cankaoxiaoxi 参考消息 pdf 1957-1998 0 likes15k downloads2y agoHugging Face19Aneeshers /tennis-sackmann-archive Tennis datasets — archive of Jeff Sackmann's data This dataset is an archival mirror of the public tennis datasets compiled by Jeff Sackmann. It exists so the data remains available and citable. It contains only data and its documentation — no models, analysis, or derived code. A matching mirror lives on GitHub: https://github.com/Aneeshers/tennis-sackmann-archive Contents Folder What it is Files Coverage slam_pointbypoint/ Point-by-point logs for the… See the full description on the dataset page: https://huggingface.co/datasets/Aneeshers/tennis-sackmann-archive.3 likes13k downloads3mo agoHugging Face20kdcyberdude /archive_pb1 likes13k downloads3y agoHugging Face21banned-historical-archives /renminribao 人民日报1946-2003 数据库+原始文件 4 likes13k downloads2y agoHugging Face22ShapeNet /ShapeNetCore-archivegatedThis repository holds archives (zip files) of main versions of ShapeNetCore, a subset of ShapeNet.ShapeNetCore is a densely annotated subset of ShapeNet covering 55 common object categories with ~51,300 unique 3D models. Each model in ShapeNetCore are linked to an appropriate synset in WordNet 3.0. Please see DATA.md for details about the data. If you use ShapeNet data, you agree to abide by the ShapeNet terms of use. You are only allowed to redistribute the data to your research associates… See the full description on the dataset page: https://huggingface.co/datasets/ShapeNet/ShapeNetCore-archive.35 likes11k downloads13d agoHugging Face23ArchEGraph /ArchEGraph ArchEGraph ArchEGraph is a building-energy dataset organized for graph-based and weather-conditioned learning. Dataset Summary Total cases in manifest.csv: 49,326 Unique buildings: 5,481 Unique weather IDs: 64 n_steps range: 968 to 8,760 n_spaces range: 1 to 231 This repository currently stores: manifest.csv (index of all cases) building/ (5,481 files) geometry/ (5,482 files) weather/ (64 files) energy/ (49,326 files; nested under subfolders like 00/) split/… See the full description on the dataset page: https://huggingface.co/datasets/ArchEGraph/ArchEGraph.tabulargraph-ml100K<n<1M1 likes10k downloads5mo agoHugging Face24OmniAICreator /ASMR-Archive-Processed ASMR-Archive-Processed (WIP) Update (2026-04-03): This dataset has reached the Hugging Face Public Storage Limit. After contacting support, we were informed that the only option is to pay for a storage expansion. Consequently, updates to this dataset are now suspended. Work in Progress — expect breaking changes while the pipeline and data layout stabilize. This dataset contains ASMR audio data sourced from DeliberatorArchiver/asmr-archive-data-01 and… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed.imageautomatic-speech-recognition96 likes10k downloads6mo agoHugging Face25banned-historical-archives /huabao-before-19492 likes10k downloads2y agoHugging Face26alexandrainst /m_arc Multilingual ARC Dataset Summary This dataset is a machine translated version of the ARC dataset. The Icelandic (is) part was translated with Miðeind's Greynir model and Norwegian (nb) was translated with DeepL. The rest of the languages was translated using GPT-3.5-turbo by the University of Oregon, and this part of the dataset was originally uploaded to this Github repository. textquestion-answering10K<n<100K4 likes9.3k downloads3y agoHugging Face27banned-historical-archives /peking-review0 likes9.1k downloads2y agoHugging Face28binhduong86224 /vineyard-vigor-vegetation-index-archive0 likes8.9k downloads1h agoHugging Face29banned-historical-archives /hangzhouribao0 likes8.5k downloads2y agoHugging Face30AiAF /SCPWiki-Cleaned-PDF-Archivesdocumenttext-generationn<1K1 likes7.7k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.