datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
banned-historical-archives
和谐历史档案馆数据集 - Banned Historical Archives Datasets
和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。
目录结构
banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步
raw # 原始文件
config # 配置文件
todo # 存放暂未录入网站的文件
部分报纸和图片资料存放在单独的仓库:
名称
地址
状态
参考消息
https://huggingface.co/datasets/banned-historical-archives/ckxx
未录入
人民日报
https://huggingface.co/datasets/banned-historical-archives/rmrb
已精选重要的文章录入
文汇报… See the full description on the dataset page: https://huggingface.co/datasets/banned-historical-archives/banned-historical-archives.ai2_arc
Dataset Card for "ai2_arc"
Dataset Summary
A new dataset of 7,787 genuine grade-school level, multiple-choice science questions, assembled to encourage research in
advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains
only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. We are also
including a corpus of over 14 million science sentences… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ai2_arc.zhongyangribaoSubmission-ArchiveAI-CUDA-Engineer-Archive
The AI CUDA Engineer Archive 👷: Agentic CUDA Kernel Discovery, Optimization & Composition
We release The AI CUDA Engineer archive, a dataset consisting of approximately 30,000 CUDA kernels generated by The AI CUDA Engineer. It is released under the CC-By-4.0 license and can be accessed via HuggingFace and interactively visualized here. The dataset is based on the Kernel tasks provided in KernelBench and includes a torch reference implementation, torch, NCU and Clang-tidy… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/AI-CUDA-Engineer-Archive.arctic
Arctic Shift Reddit Archive
Every Reddit comment and submission since 2005, organized as monthly Parquet shards
What is it?
The full Reddit archive from Arctic Shift, converted to Parquet and hosted here for easy access. Covers every public subreddit from 2005-12 through 2026-02.
Right now the archive has 15.7B items (12.9B comments, 2.8B submissions) in 1.3 TB of compressed Parquet. Comments and submissions are stored as separate datasets, split into monthly… See the full description on the dataset page: https://huggingface.co/datasets/open-index/arctic.picbreeder-vlm-archive
Picbreeder-VLM Archive
Every image evolved by the swarm of vision-language-model "breeders" in
In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models
(GECCO 2026), together with the CPPN genomes that produced them, the agents' reasoning transcripts, the
lineage graphs, and the analysis artifacts behind the paper and blog.
The original Picbreeder (Secretan et al., 2008) let crowds of
humans collaboratively evolve images from
CPPN… See the full description on the dataset page: https://huggingface.co/datasets/picbreeder-vlm/picbreeder-vlm-archive.State-Parse-FilteredThe single cell RNA-seq dataset with human PBMC samples was sourced from Parse Biosciences [1]. [1] Performance of Evercode™ WT v3 in Human Immune Cells (PBMCs), https://www.parsebiosciences.com/datasets/performance-of-evercode-wt-v3-in-human-immune-cells-pbmcs/; Parse Biosciences, Seattle, USA; accessed 05/27/2025.
Certain uses of this data may require a license from Parse Biosciences, Inc.
jfk-archives
Dataset Card for JFK Archives
This dataset is a collection of all records pertaining to the assassination of the
US president, John F. Kennedy, released until April 2025 through archives.org
by the US government.
Dataset Details
Dataset Description
The original data downloaded from archives.org
consists of 56,300 scanned documents in PDF format, released until April 2025. The files are
organized by their release year(s): 2107-2018, 2021, 2022, 2023 and 2025.… See the full description on the dataset page: https://huggingface.co/datasets/farhanhubble/jfk-archives.opengenome2
OpenGenome2
OpenGenome2 is a database of nearly 9 trillion base pairs of curated DNA from across all domains of life. Collected from diverse species and public data sources, OpenGenome2 was used to train Evo 2 models. Please refer to the Evo 2 preprint or github repository for further details and usage examples.
We provide OpenGenome2 in two formats, the dataset is organized into two main directories to reflect this:
fasta which contain the DNA sequences
jsonl which… See the full description on the dataset page: https://huggingface.co/datasets/arcinstitute/opengenome2.low-prioritydistilabel-intel-orca-dpo-pairs-binarizedThis is the binarized version of distilabel Orca Pairs for DPO and ORPO.
Reference: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs?row=0
arc-whestbench-public-2026
Organized by:
Alignment Research Center (ARC),
AIcrowd
WhestBench 2026: ARC White-Box Estimation Challenge
WhestBench is a benchmark for white-box activation estimation: given the weights of a randomly initialized ReLU multi-layer perceptron (MLP) and a strict floating-point-operation (FLOP) budget, predict the average post-activation value of every neuron when the network is fed standard Gaussian inputs.
This is the WhestBench 2026… See the full description on the dataset page: https://huggingface.co/datasets/aicrowd/arc-whestbench-public-2026.cmu-arctic-xvectors
Speaker embeddings extracted from CMU ARCTIC
There is one .npy file for each utterance in the dataset, 7931 files in total. The speaker embeddings are 512-element X-vectors.
The CMU ARCTIC dataset divides the utterances among the following speakers:
bdl (US male)
slt (US female)
jmk (Canadian male)
awb (Scottish male)
rms (US male)
clb (US female)
ksp (Indian male)
The X-vectors were extracted using this script, which uses the speechbrain/spkrec-xvect-voxceleb model.
Usage:
from… See the full description on the dataset page: https://huggingface.co/datasets/Matthijs/cmu-arctic-xvectors.Stack-scBaseCount189Mdim-discovery-archive
Geometry of Decision Making in Language Models
Abhinav Joshi · Divyanshu Bhatt · Ashutosh ModiNeurIPS 2025
This repository contains the official implementation/release for the NeurIPS 2025 paper Geometry of Decision Making in Language Models.
We study the internal decision-making processes of large language models through the lens of intrinsic dimension (ID), analyzing how hidden representations evolve across layers in a multiple-choice… See the full description on the dataset page: https://huggingface.co/datasets/Exploration-Lab/dim-discovery-archive.banned-historical-archives
和谐历史档案馆数据集 - Banned Historical Archives Datasets
和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。
目录结构
banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步
raw # 原始文件
config # 配置文件
todo # 存放暂未录入网站的文件
部分报纸和图片资料存放在单独的仓库:
名称
地址
状态
参考消息
https://huggingface.co/datasets/banned-historical-archives/ckxx
未录入
人民日报
https://huggingface.co/datasets/banned-historical-archives/rmrb
已精选重要的文章录入
文汇报… See the full description on the dataset page: https://huggingface.co/datasets/Dragonegg2026/banned-historical-archives.cankaoxiaoxi
参考消息 pdf 1957-1998
tennis-sackmann-archive
Tennis datasets — archive of Jeff Sackmann's data
This dataset is an archival mirror of the public tennis datasets compiled by Jeff Sackmann.
It exists so the data remains available and citable. It contains only data and its documentation —
no models, analysis, or derived code.
A matching mirror lives on GitHub: https://github.com/Aneeshers/tennis-sackmann-archive
Contents
Folder
What it is
Files
Coverage
slam_pointbypoint/
Point-by-point logs for the… See the full description on the dataset page: https://huggingface.co/datasets/Aneeshers/tennis-sackmann-archive.archive_pbrenminribao
人民日报1946-2003 数据库+原始文件
ShapeNetCore-archiveThis repository holds archives (zip files) of main versions of ShapeNetCore, a subset of ShapeNet.ShapeNetCore is a densely annotated subset of ShapeNet covering 55 common object categories with ~51,300 unique 3D models. Each model in ShapeNetCore are linked to an appropriate synset in WordNet 3.0.
Please see DATA.md for details about the data.
If you use ShapeNet data, you agree to abide by the ShapeNet terms of use. You are only allowed to redistribute the data to your research associates… See the full description on the dataset page: https://huggingface.co/datasets/ShapeNet/ShapeNetCore-archive.ArchEGraph
ArchEGraph
ArchEGraph is a building-energy dataset organized for graph-based and weather-conditioned learning.
Dataset Summary
Total cases in manifest.csv: 49,326
Unique buildings: 5,481
Unique weather IDs: 64
n_steps range: 968 to 8,760
n_spaces range: 1 to 231
This repository currently stores:
manifest.csv (index of all cases)
building/ (5,481 files)
geometry/ (5,482 files)
weather/ (64 files)
energy/ (49,326 files; nested under subfolders like 00/)
split/… See the full description on the dataset page: https://huggingface.co/datasets/ArchEGraph/ArchEGraph.ASMR-Archive-Processed
ASMR-Archive-Processed (WIP)
Update (2026-04-03): This dataset has reached the Hugging Face Public Storage Limit. After contacting support, we were informed that the only option is to pay for a storage expansion. Consequently, updates to this dataset are now suspended.
Work in Progress — expect breaking changes while the pipeline and data layout stabilize.
This dataset contains ASMR audio data sourced from DeliberatorArchiver/asmr-archive-data-01 and… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed.huabao-before-1949m_arc
Multilingual ARC
Dataset Summary
This dataset is a machine translated version of the ARC dataset.
The Icelandic (is) part was translated with Miðeind's Greynir model and Norwegian (nb) was translated with DeepL. The rest of the languages was translated using GPT-3.5-turbo by the University of Oregon, and this part of the dataset was originally uploaded to this Github repository.
peking-reviewvineyard-vigor-vegetation-index-archivehangzhouribaoSCPWiki-Cleaned-PDF-Archives
