datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
banned-historical-archives
和谐历史档案馆数据集 - Banned Historical Archives Datasets
和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。
目录结构
banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步
raw # 原始文件
config # 配置文件
todo # 存放暂未录入网站的文件
部分报纸和图片资料存放在单独的仓库:
名称
地址
状态
参考消息
https://huggingface.co/datasets/banned-historical-archives/ckxx
未录入
人民日报
https://huggingface.co/datasets/banned-historical-archives/rmrb
已精选重要的文章录入
文汇报… See the full description on the dataset page: https://huggingface.co/datasets/banned-historical-archives/banned-historical-archives.zhongyangribaoSubmission-ArchiveAI-CUDA-Engineer-Archive
The AI CUDA Engineer Archive 👷: Agentic CUDA Kernel Discovery, Optimization & Composition
We release The AI CUDA Engineer archive, a dataset consisting of approximately 30,000 CUDA kernels generated by The AI CUDA Engineer. It is released under the CC-By-4.0 license and can be accessed via HuggingFace and interactively visualized here. The dataset is based on the Kernel tasks provided in KernelBench and includes a torch reference implementation, torch, NCU and Clang-tidy… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/AI-CUDA-Engineer-Archive.picbreeder-vlm-archive
Picbreeder-VLM Archive
Every image evolved by the swarm of vision-language-model "breeders" in
In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models
(GECCO 2026), together with the CPPN genomes that produced them, the agents' reasoning transcripts, the
lineage graphs, and the analysis artifacts behind the paper and blog.
The original Picbreeder (Secretan et al., 2008) let crowds of
humans collaboratively evolve images from
CPPN… See the full description on the dataset page: https://huggingface.co/datasets/picbreeder-vlm/picbreeder-vlm-archive.jfk-archives
Dataset Card for JFK Archives
This dataset is a collection of all records pertaining to the assassination of the
US president, John F. Kennedy, released until April 2025 through archives.org
by the US government.
Dataset Details
Dataset Description
The original data downloaded from archives.org
consists of 56,300 scanned documents in PDF format, released until April 2025. The files are
organized by their release year(s): 2107-2018, 2021, 2022, 2023 and 2025.… See the full description on the dataset page: https://huggingface.co/datasets/farhanhubble/jfk-archives.low-prioritydim-discovery-archive
Geometry of Decision Making in Language Models
Abhinav Joshi · Divyanshu Bhatt · Ashutosh ModiNeurIPS 2025
This repository contains the official implementation/release for the NeurIPS 2025 paper Geometry of Decision Making in Language Models.
We study the internal decision-making processes of large language models through the lens of intrinsic dimension (ID), analyzing how hidden representations evolve across layers in a multiple-choice… See the full description on the dataset page: https://huggingface.co/datasets/Exploration-Lab/dim-discovery-archive.banned-historical-archives
和谐历史档案馆数据集 - Banned Historical Archives Datasets
和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。
目录结构
banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步
raw # 原始文件
config # 配置文件
todo # 存放暂未录入网站的文件
部分报纸和图片资料存放在单独的仓库:
名称
地址
状态
参考消息
https://huggingface.co/datasets/banned-historical-archives/ckxx
未录入
人民日报
https://huggingface.co/datasets/banned-historical-archives/rmrb
已精选重要的文章录入
文汇报… See the full description on the dataset page: https://huggingface.co/datasets/Dragonegg2026/banned-historical-archives.cankaoxiaoxi
参考消息 pdf 1957-1998
tennis-sackmann-archive
Tennis datasets — archive of Jeff Sackmann's data
This dataset is an archival mirror of the public tennis datasets compiled by Jeff Sackmann.
It exists so the data remains available and citable. It contains only data and its documentation —
no models, analysis, or derived code.
A matching mirror lives on GitHub: https://github.com/Aneeshers/tennis-sackmann-archive
Contents
Folder
What it is
Files
Coverage
slam_pointbypoint/
Point-by-point logs for the… See the full description on the dataset page: https://huggingface.co/datasets/Aneeshers/tennis-sackmann-archive.archive_pbrenminribao
人民日报1946-2003 数据库+原始文件
ShapeNetCore-archiveThis repository holds archives (zip files) of main versions of ShapeNetCore, a subset of ShapeNet.ShapeNetCore is a densely annotated subset of ShapeNet covering 55 common object categories with ~51,300 unique 3D models. Each model in ShapeNetCore are linked to an appropriate synset in WordNet 3.0.
Please see DATA.md for details about the data.
If you use ShapeNet data, you agree to abide by the ShapeNet terms of use. You are only allowed to redistribute the data to your research associates… See the full description on the dataset page: https://huggingface.co/datasets/ShapeNet/ShapeNetCore-archive.ASMR-Archive-Processed
ASMR-Archive-Processed (WIP)
Update (2026-04-03): This dataset has reached the Hugging Face Public Storage Limit. After contacting support, we were informed that the only option is to pay for a storage expansion. Consequently, updates to this dataset are now suspended.
Work in Progress — expect breaking changes while the pipeline and data layout stabilize.
This dataset contains ASMR audio data sourced from DeliberatorArchiver/asmr-archive-data-01 and… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed.huabao-before-1949peking-reviewvineyard-vigor-vegetation-index-archivehangzhouribaoSCPWiki-Cleaned-PDF-ArchivesscheduleSee https://github.com/ust-archive/ust-archive for more information.
rl-run-archive-2026
RL run archive 2026
Archived raw run artifacts (rollout trajectories, rendered frames, policy and optimizer
checkpoints, configs, logs) from simulation reinforcement-learning experiments, published for
long-term preservation and reproducibility.
Layout mirrors the verified backup trees they were copied from:
tilde/20260915-102000/ and taurus/20260915-085631/: batched tar archives. Every archive
carries a per-file SHA-256 manifest inside it; the batch inventories (9998.json.gz… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/rl-run-archive-2026.SCPWiki-Archive-02-March-2025-Datasetsbanned-historical-archives
和谐历史档案馆数据集 - Banned Historical Archives Datasets
和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。
目录结构
banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步
raw # 原始文件
config # 配置文件
todo # 存放暂未录入网站的文件
部分报纸和图片资料存放在单独的仓库:
名称
地址
状态
参考消息
https://huggingface.co/datasets/banned-historical-archives/ckxx
未录入
人民日报
https://huggingface.co/datasets/banned-historical-archives/rmrb
已精选重要的文章录入
文汇报… See the full description on the dataset page: https://huggingface.co/datasets/RonaldoDD/banned-historical-archives.CEF_Main_Archive
Core Emotion Framework (CEF) Main Archive
The Decalogue of Operators
The Core Emotion Framework defines exactly ten functional operators. This is the complete and authoritative set. No additional operators exist. No operators may be removed, renamed, or substituted. This dataset serves as the absolute source of truth for the following:
Sensing
Calculating
Deciding
Expanding
Constricting
Achieving
Arranging
Appreciating
Boosting
Accepting
{
"@context":… See the full description on the dataset page: https://huggingface.co/datasets/CoreEmotionFramework/CEF_Main_Archive.codetorch-archivestream-archive
stream-archive
Twitch and Kick chat logs from Italian streamers.
Streamers
Twitch.com (39)
aladinottv, alisonrevenge, arkanightlive, bartopanzer, billybella_, dankol83, dariomocciatwitch, davidrubino, diariodelrusso, enkk, federicacasula_, fufflix, grenbaud, gskianto, homyatol, ilgabbrone, ilrossopiubelloditwitch, immortale____, kasumisen, lollolacustre, lucakingm, luiskant690, macchiativincenzo_babbohs, marcomerrino, menestointhailandia… See the full description on the dataset page: https://huggingface.co/datasets/deplana/stream-archive.archive-dolma3-mix-150b-enriched
archive-dolma3-mix-150b-enriched
ARCHIVE (pre-6T era): WebOrganizer-enriched variant of the 150B Dolma3 mix (not the pool).
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_mix_150B_enriched
Renamed
2026-05-25
See docs/data_home/inventory.json for the full inventory… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-mix-150b-enriched.shorouk-pdf-archive
Shorouk PDF archive
Original newspaper PDFs from Al Shorouk's public archive. Filenames use YYYY-MM-DD.pdf.
As verified in this update, the dataset contains 6,276 PDFs, dated 2009-02-01 through 2026-09-08. The September 8 edition was already available from the publisher when checked on September 7 in New York.
Coverage and remaining gaps
This update added 518 issues and filled all calendar gaps from January 1, 2025 through September 8, 2026. 153 earlier dates… See the full description on the dataset page: https://huggingface.co/datasets/eltokh7/shorouk-pdf-archive.free-music-archive-full
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-full.
