datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arctic
Arctic Shift Reddit Archive
Every Reddit comment and submission since 2005, organized as monthly Parquet shards
What is it?
The full Reddit archive from Arctic Shift, converted to Parquet and hosted here for easy access. Covers every public subreddit from 2005-12 through 2026-02.
Right now the archive has 1.6B items (362.1M comments, 1.2B submissions) in 181.4 GB of compressed Parquet. Comments and submissions are stored as separate datasets, split into monthly shards… See the full description on the dataset page: https://huggingface.co/datasets/Dk587/arctic.openvino-arc140v-lunarlake
OpenVINO local-inference on an Intel Arc 140V (Lunar Lake) iGPU
Reference performance data for running local models on a single Intel Core Ultra 7 258V
(Lunar Lake) laptop with the integrated Intel Arc 140V (Xe2) GPU, via OpenVINO. All
inference runs on the iGPU; the NPU stays idle throughout, confirmed by the telemetry here.
This is reference characterization shared by a non-expert contributor — careful measurements
on one machine, offered so others can compare and correct, not… See the full description on the dataset page: https://huggingface.co/datasets/blairducrayoppat/openvino-arc140v-lunarlake.Technical-Architectures-Large
Technical Architectures Large (294k Samples)
Overview
Generating complex, syntactically valid diagram code from natural language requirements is a major challenge for AI models. This dataset bridges that gap by providing over 293,000+ distinct enterprise software architectures generated using two cutting-edge models: GPT-OSS-120B and Qwen3-Coder-Next-FP8.
Unlike simple "toy" examples, these architectures model realistic enterprise systems complete with client… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Technical-Architectures-Large.jam-rollout-arc-evals
Rollout arc — raw generations
Every model generation behind the write-ups in
mcp-tool-shop-org/ai-jam-sessions
under experiments/rollout-arc/p4/.
Two things you can do with this.
Check our arithmetic. The repo has the readout scripts, the preregistrations and the
intervals — but the generations they were computed from are ~51 MB and were never committed, so
a clone got the conclusions and no way to recompute them. These are those files, unfiltered.
Or run the loop yourself. The… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-rollout-arc-evals.Rail_Freight_Logistics_Company_Email_Archive_Sample
Ukrainian Rail-Freight Correspondence Corpus (Sample)
Real operational correspondence from a working freight forwarding business, and the
documents attached to it — consignment notes, service acts, invoices, wagon
manifests. Not scraped, not synthetic, and never published anywhere before.
This is a de-identified sample released for evaluation. It is drawn from a larger
private archive; see Full archive below.
Published by Akuma London · akumalondon.com
Why this… See the full description on the dataset page: https://huggingface.co/datasets/akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample.sayit-archive-tw
Audrey Tang Transcript Corpus
Public transcripts of Audrey Tang (唐鳳) — 2025 Right Livelihood Laureate, civic hacker, and Taiwan's Cyber Ambassador.
Tang is co-author of Plurality: The Future of Collaborative Technology and Democracy and an inaugural Senior Accelerator Fellow at the Oxford Institute for Ethics in AI. She served as Taiwan's first Digital Minister (2016–2024) and the world's first nonbinary cabinet minister, awarded the Right Livelihood Award for "advancing the social… See the full description on the dataset page: https://huggingface.co/datasets/audreyt/sayit-archive-tw.discord-archive
Discord Archive
This is an archive of messages from the Banodoco Discord community, where
technical and artistic practitioners have been discussing open source AI art for
the past three years.
The archive captures a long-running community record of people learning,
training, evaluating, and using open source AI art models in practice. It
contains discussion around model releases, workflows, tooling, troubleshooting,
creative experiments, training details, and the many small… See the full description on the dataset page: https://huggingface.co/datasets/Banodoco/discord-archive.review_arcade
ACL ARR Reviews - Review Arcade Project
This dataset contains paper reviews from the ACL ARR (Association for Computational Linguistics - Annual Review of Research) program. The reviews are organized into splits corresponding to papers that were accepted or rejected for publication, as well as specific subsets used for the research analysis.
This dataset was introduced in the paper Review Arcade: On the Human Alignment and Gameability of LLM Reviews.
Code:… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/review_arcade.arc-agi-2-grids
ARC-AGI-2 Grids — training + analysis corpus (NVARC-compatible)
Companion dataset for the Kaggle ARC Prize 2026 (ARC-AGI-2) solver built on
sorokin/qwen3_4b_grids15_sft139 + per-task rank-256 LoRA (NVARC lineage).
Everything here is generated from public canonical data only (1,000
training / 120 evaluation tasks); no hidden competition data is included.
Contents
Path
Rows
Description
train/train_tasks.jsonl
1,000
canonical training tasks (full I/O)… See the full description on the dataset page: https://huggingface.co/datasets/Nabidnur/arc-agi-2-grids.agent-archive
Agent Archive: PR Interactions
This dataset contains pull request interactions from GitHub repositories where AI coding agents (Claude Code, Copilot, etc.) contributed code. The data was sourced from the AgentPack dataset and enriched with full PR timeline events from the GitHub API.
Dataset Description
Each row represents a single pull request with its complete interaction history, including comments, reviews, commits, and other timeline events.
Column… See the full description on the dataset page: https://huggingface.co/datasets/nuprl-staging/agent-archive.Technical-Architectures-Large
Technical Architectures Large (210k+ Samples)
Overview
Generating complex, syntactically valid diagram code from natural language requirements is a major challenge for AI models. This dataset bridges that gap by providing over 210,000 distinct enterprise software architectures generated using two cutting-edge models: GPT-OSS-120B and Qwen3-Coder-Next-FP8.
Unlike simple "toy" examples, these architectures model realistic enterprise systems complete with client… See the full description on the dataset page: https://huggingface.co/datasets/autoshift/Technical-Architectures-Large.arc-agi-augmented-100
ARC-AGI Augmented Dataset
This dataset is an augmented version of the Abstraction and Reasoning Corpus (ARC-AGI), processed for training neural networks (such as Transformers or Neural Cellular Automata).
Dataset Details
Original Source: ARC-AGI Benchmark
License: MIT
Augmentation Method:
Dihedral Transformations: 8 symmetries (rotations/flips).
Color Permutation: Random permutation of colors 1-9 (0 is fixed as background).
Translational Padding: Randomly positioning the… See the full description on the dataset page: https://huggingface.co/datasets/KotshinZ/arc-agi-augmented-100.code-deontologie-architectes
Code de déontologie des architectes, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-deontologie-architectes.bell-labs-technical-archive
Bell Labs Documents and Stuff
This is a conservative public-release subset of the internal BELLA continued-pretraining corpus. It keeps the Bell-system technical material that survived a stricter final pass for public dataset hosting and removes records that still looked risky, off-scope, or too low-signal for a Hugging Face corpus listing.
What is in the release
Split
Documents
train
1220
validation
29
test
42
The release contains 1291 documents out… See the full description on the dataset page: https://huggingface.co/datasets/hunterbown/bell-labs-technical-archive.atcoder_arc_contests
Notification
Atcoder is selling this data now. If you are interested in accessing it please contact them.
Dataset Summary
This dataset aims to facilitate the creation of sophisticated, multi-turn dialogue datasets focused on coding
that could be used for training reasoning Large Language Models (LLMs), particularly for Supervised Fine-Tuning (SFT) and Knowledge Distillation techniques.
It also serves as a robust foundation for problem-solving in Large Language… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/atcoder_arc_contests.neural-architecture-search-papers
Neural Architecture Search (NAS) Papers — FineSet
A research-paper dataset on Neural Architecture Search (NAS) Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on Neural Architecture Search (NAS) Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/neural-architecture-search-papers.
