datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
banned-historical-archives
和谐历史档案馆数据集 - Banned Historical Archives Datasets
和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。
目录结构
banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步
raw # 原始文件
config # 配置文件
todo # 存放暂未录入网站的文件
部分报纸和图片资料存放在单独的仓库:
名称
地址
状态
参考消息
https://huggingface.co/datasets/banned-historical-archives/ckxx
未录入
人民日报
https://huggingface.co/datasets/banned-historical-archives/rmrb
已精选重要的文章录入
文汇报… See the full description on the dataset page: https://huggingface.co/datasets/banned-historical-archives/banned-historical-archives.picbreeder-vlm-archive
Picbreeder-VLM Archive
Every image evolved by the swarm of vision-language-model "breeders" in
In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models
(GECCO 2026), together with the CPPN genomes that produced them, the agents' reasoning transcripts, the
lineage graphs, and the analysis artifacts behind the paper and blog.
The original Picbreeder (Secretan et al., 2008) let crowds of
humans collaboratively evolve images from
CPPN… See the full description on the dataset page: https://huggingface.co/datasets/picbreeder-vlm/picbreeder-vlm-archive.banned-historical-archives
和谐历史档案馆数据集 - Banned Historical Archives Datasets
和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。
目录结构
banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步
raw # 原始文件
config # 配置文件
todo # 存放暂未录入网站的文件
部分报纸和图片资料存放在单独的仓库:
名称
地址
状态
参考消息
https://huggingface.co/datasets/banned-historical-archives/ckxx
未录入
人民日报
https://huggingface.co/datasets/banned-historical-archives/rmrb
已精选重要的文章录入
文汇报… See the full description on the dataset page: https://huggingface.co/datasets/Dragonegg2026/banned-historical-archives.ASMR-Archive-Processed
ASMR-Archive-Processed (WIP)
Update (2026-04-03): This dataset has reached the Hugging Face Public Storage Limit. After contacting support, we were informed that the only option is to pay for a storage expansion. Consequently, updates to this dataset are now suspended.
Work in Progress — expect breaking changes while the pipeline and data layout stabilize.
This dataset contains ASMR audio data sourced from DeliberatorArchiver/asmr-archive-data-01 and… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed.ultimateArchWorldfamily-archival-scans
Family archival scans
Page scans of primary-source archival records used in the bobpanil/family
genealogy project. Public and world-readable, so tools and agents without
credentials can fetch pages directly over plain HTTPS.
6,475 files across 50 archival units from five holding institutions, plus a
database-index folder and a screenshot folder. All records are pre-1943 and
concern people long deceased; no information about living individuals is
included.
Structure… See the full description on the dataset page: https://huggingface.co/datasets/bobpanil/family-archival-scans.meta-archiveeastbrush_archive
Eastbrush Archive
Official Website (Full Archive System):
https://www.eastbrush.com
This dataset contains high-resolution images and structured tags for AI training.
The full archive system — including chapter exhibitions, structural context, and extended records — is available on the official website.
What is my true self?
The Eastbrush Archive is a long-term, evolving system that documents the visual language ofJang Byeong Eun (Eastbrush / 張炳彥) — a painter whose… See the full description on the dataset page: https://huggingface.co/datasets/eastbrush/eastbrush_archive.scratch-archivecs2-demo-archive
CS2 Demo to Dataset — Pipeline Output Samples
Sample archives produced by the open-source
cs2-demo-to-dataset
pipeline, which converts a single CS2 .dem replay file into per-round,
per-player first-person video aligned to tick-level state, input and event
tables.
This release is not a dataset contribution. The point of the upload is to
demonstrate that the pipeline produces a coherent, reproducible archive
format. Please see the
GitHub repository for the
recorder code… See the full description on the dataset page: https://huggingface.co/datasets/Vasy7777/cs2-demo-archive.cbi-archive-raw
Central Bank of Ireland Archive: original source files
6,309 original files, 6.56 GB. Every PDF, spreadsheet, Word document and
archive gathered from the Central Bank of Ireland's public website, stored by
content hash so that a search result can be turned back into the document a
human would actually read.
This is the raw tier. If you want the text, you almost certainly want
aditya487/cbi-archive-corpus
instead: 5,568 documents and 89,242 page or pseudo-page rows as Parquet… See the full description on the dataset page: https://huggingface.co/datasets/aditya487/cbi-archive-raw.ArchDailyhkgongshangwanbaocorpus-archive
corpus-archive
[!WARNING]
Experimental Dataset Architecture: The repository structure, metadata tiers, category taxonomies, and catalog indexing formats are currently under active design and evaluation. All specifications, metadata keys, and JSON schemas detailed below represent representational examples and intended targets.
This repository serves as a structured digital textual archive preserving Hmar literature, historical accounts, school textbooks, dictionaries, parallel… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/corpus-archive.hkgongshangribaodagongbaomodern-architecturehuaqiaoribaodanbooru2023
[Mirror]Danbooru2023: A Large-Scale Crowdsourced and Tagged Anime Illustration Dataset
Danbooru2023 is an extension of Danbooru2021, featuring over 6.8 million anime-style images, totaling more than 8.3 TB.
Each image is accompanied by community-contributed tags that provide detailed descriptions of its content, including characters,
artists, copyright information, concepts, and attire.
This makes it a crucial resource for stylized computer vision tasks and transfer learning.… See the full description on the dataset page: https://huggingface.co/datasets/zenless-archive/danbooru2023.AraMS-Restore
AraMS-Restore — Real Damaged Arabic Manuscript Lines
177 line images cropped from real damaged pages of a historical Arabic
manuscript (book_09), each with its transcription. This is the evaluation
input for AraMS-Restore: the
restoration models are trained on synthetic degradation, and these lines are the
honest test of whether that transfers to genuine manuscript decay.
There are no clean counterparts and no ground-truth restored images — the damage
is what was on the page.… See the full description on the dataset page: https://huggingface.co/datasets/Archatext/AraMS-Restore.architectsstride-architecture-components-v1
STRIDE Architecture Threat Modeling Dataset (AWS & Azure)
📌 Overview
This dataset was created to enable automatic STRIDE threat modeling from cloud architecture diagrams (AWS and Azure).
The goal is to detect architectural components in diagrams and support automated threat identification based on data flows and trust boundaries.
Annotations were created using Label Studio in YOLO format.
Total images: 4190Total classes: 32
🎯 Purpose
Detect cloud… See the full description on the dataset page: https://huggingface.co/datasets/guillherms/stride-architecture-components-v1.arch_gastric
Dataset Card for "arch_gastric"
More Information needed
nanochat
nanochat
nanochat is the simplest experimental harness for training LLMs. It is designed to run on a single GPU node, the code is minimal/hackable, and it covers all major LLM stages including tokenization, pretraining, finetuning, evaluation, inference, and a chat UI. For example, you can train your own GPT-2 capability LLM (which cost $43,000 to train in 2019) for only $48 (2 hours of 8XH100 GPU node) and then talk to it in a familiar ChatGPT-like web UI. On a spot instance… See the full description on the dataset page: https://huggingface.co/datasets/ArchaeonSeq/nanochat.arch-building-dataset
World Architectural Buildings Dataset (FGIC) for Multi‑Class Image Classification
Multi‑Class Image Classification dataset of world architectural buildings with finalized curation.
Classes
Class
Count
Description
barn
1,680
Traditional wooden barn architecture — residential and storage buildings
bridge
1,680
Various bridge architectures (suspension, arch, truss)
castle
1,680
Medieval and modern castle structures
mosque
1,680
Islamic mosque… See the full description on the dataset page: https://huggingface.co/datasets/0xgr3y/arch-building-dataset.basil-instances-archive-3
maximilian-franz/basil-instances
Per-plant-instance segmented crops derived from ['maximilian-franz/basil', 'maximilian-franz/basil-2'], one row per
(original frame × confirmed plant instance).
Layout
ImageFolder-style dataset: metadata.csv at the repo root, with a file_name column pointing
to each row's masked crop under images/<plant_instance_id>/<NNNN>.png, and a
bbox_file_name column pointing to the same row's plain rectangular crop under… See the full description on the dataset page: https://huggingface.co/datasets/maximilian-franz/basil-instances-archive-3.basil-instances-archive-2
maximilian-franz/basil-instances
Per-plant-instance segmented crops derived from ['maximilian-franz/basil', 'maximilian-franz/basil-2'], one row per
(original frame × confirmed plant instance).
Layout
ImageFolder-style dataset: metadata.csv at the repo root, with a file_name column pointing
to each row's image under images/<plant_instance_id>/<NNNN>.png. Load with:
from datasets import load_dataset
ds = load_dataset("maximilian-franz/basil-instances")… See the full description on the dataset page: https://huggingface.co/datasets/maximilian-franz/basil-instances-archive-2.razavi-benchRazavi-bench
An expert-curated benchmark for analog-design reasoning.
Razavi-bench packages the question-answer assessments from Behzad Razavi's
Analog Design Experiments With AI Part 1 and Part 2 into a clean
one-task-per-directory benchmark. The tasks probe whether a model can reason
about MOS devices, small-signal circuits, feedback, oscillators, comparators,
dividers, LNAs, TIAs, and LC oscillators.
Each task directory keeps only the benchmark prompt, figure, and curated… See the full description on the dataset page: https://huggingface.co/datasets/Arcadia-2026/razavi-bench.SoftVTBench-archive
SoftVTBench — archive
Frozen snapshot of everything that lived in Arthur12137/SoftVTBench before the
2026-08-13 re-release: the evaluation USD assets, the soft-body assets, and the
first partial data drops.
This repo is not maintained. The current dataset is at
Arthur12137/SoftVTBench.
Contents: eval-assets/, soft-assets/, object-rigid/, object-soft/,
spatial-rigid/, spatial-soft/ (9319 files, 2.3 GB).
Original dataset card (kept verbatim)
SoftVTBench dataset… See the full description on the dataset page: https://huggingface.co/datasets/Arthur12137/SoftVTBench-archive.
