datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ArchEGraph
ArchEGraph
ArchEGraph is a building-energy dataset organized for graph-based and weather-conditioned learning.
Dataset Summary
Total cases in manifest.csv: 49,326
Unique buildings: 5,481
Unique weather IDs: 64
n_steps range: 968 to 8,760
n_spaces range: 1 to 231
This repository currently stores:
manifest.csv (index of all cases)
building/ (5,481 files)
geometry/ (5,482 files)
weather/ (64 files)
energy/ (49,326 files; nested under subfolders like 00/)
split/… See the full description on the dataset page: https://huggingface.co/datasets/ArchEGraph/ArchEGraph.arc-agi-3-schema-traces
ARC-AGI-3 Schema Gameplay Trajectories
This release contains 50 ARC-AGI-3 gameplay trajectories and a dependency-free
scoring utility. The trajectories are split evenly across two collections:
gpt_5_6_sol/: 25 GPT-5.6 Sol trajectories.
claude_fable_opus/: 25 trajectories from Claude Opus 4.8 and Claude Fable 5.
Each trajectory directory includes run.json, a streamed events.jsonl event
log, sanitized session data, snapshots, and the shareable text/image files
produced during… See the full description on the dataset page: https://huggingface.co/datasets/schema-harness/arc-agi-3-schema-traces.cbi-archive-raw
Central Bank of Ireland Archive: original source files
6,309 original files, 6.56 GB. Every PDF, spreadsheet, Word document and
archive gathered from the Central Bank of Ireland's public website, stored by
content hash so that a search result can be turned back into the document a
human would actually read.
This is the raw tier. If you want the text, you almost certainly want
aditya487/cbi-archive-corpus
instead: 5,568 documents and 89,242 page or pseudo-page rows as Parquet… See the full description on the dataset page: https://huggingface.co/datasets/aditya487/cbi-archive-raw.arct
The Argument Reasoning Comprehension Task: Identification and Reconstruction of Implicit Warrants
https://github.com/UKPLab/argument-reasoning-comprehension-task
@InProceedings{Habernal.et.al.2018.NAACL.ARCT,
title = {The Argument Reasoning Comprehension Task: Identification
and Reconstruction of Implicit Warrants},
author = {Habernal, Ivan and Wachsmuth, Henning and
Gurevych, Iryna and Stein, Benno},
publisher = {Association for… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/arct.w2t-llm-arc-easy-lora
W2T Llm Arc Easy Lora
This repository contains artifacts for the W2T paper:
Paper: W2T: LoRA Weights Already Know What They Can Do
Repo: Weight2Token
Summary
ARC-Easy LoRA checkpoints and prepared metadata used for performance prediction.
Source Status
Storage location: local
Verification status: confirmed
Files
See manifest.json for the exact local or remote source paths used to prepare this release.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/Xiaolong-Han/w2t-llm-arc-easy-lora.us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites
Historical US layoffs archive: 6,799 WARN Act notices that state websites no longer list (2000-2025), recovered
Rebuilt 2026-09-23. Five state labor agencies — Connecticut, Michigan, New York,
North Carolina and Pennsylvania — retired the web pages their older WARN Act
layoff notices lived on. Their current pages start years later. This dataset is
every notice in our file that came from one of those retired pages and is not
on the agency's live page today: 6,799 notices, 6,799… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites.openvino-arc140v-lunarlake
OpenVINO local-inference on an Intel Arc 140V (Lunar Lake) iGPU
Reference performance data for running local models on a single Intel Core Ultra 7 258V
(Lunar Lake) laptop with the integrated Intel Arc 140V (Xe2) GPU, via OpenVINO. All
inference runs on the iGPU; the NPU stays idle throughout, confirmed by the telemetry here.
This is reference characterization shared by a non-expert contributor — careful measurements
on one machine, offered so others can compare and correct, not… See the full description on the dataset page: https://huggingface.co/datasets/blairducrayoppat/openvino-arc140v-lunarlake.ArcBench
ArcBench: ML Conference Oral Paper-Presentation Benchmark
This benchmark is from the paper Narrative-Driven Paper-to-Slide Generation via ArcDeck.
A curated benchmark dataset of 100 oral presentation paper-slide deck link pairs from top-tier machine learning conferences (CVPR, ICCV, ICLR, ICML, NeurIPS), spanning 2022–2025. Each entry provides rich metadata together with links to the original paper PDF and presentation slides, plus a script that downloads them all in one step.… See the full description on the dataset page: https://huggingface.co/datasets/ArcDeck/ArcBench.arc-agi-3-schema-traces
ARC-AGI-3 Schema Gameplay Trajectories
This release contains 50 ARC-AGI-3 gameplay trajectories and a dependency-free
scoring utility. The trajectories are split evenly across two collections:
gpt_5_6_sol/: 25 GPT-5.6 Sol trajectories.
claude_fable_opus/: 25 trajectories from Claude Opus 4.8 and Claude Fable 5.
Each trajectory directory includes run.json, a streamed events.jsonl event
log, sanitized session data, snapshots, and the shareable text/image files
produced during… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/arc-agi-3-schema-traces.2025_Virtual_Cell_Challenge_Test_DataARCHIVE-TEXT-URLS
Internet Archive English Text URLs Dataset
Dataset Description
This dataset contains 11,151,637 direct download URLs to OCR-processed text files from the Internet Archive's digital library. All entries are English-language texts spanning books, documents, historical records, and various other written materials.
Dataset Summary
Total Rows: 11,151,637
Language: English
Source: Internet Archive
Format: CSV with metadata and direct text file URLs
Text… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/ARCHIVE-TEXT-URLS.ArchEGraph-demo
ArchEGraph-demo
ArchEGraph-demo is a compact demo package of the ArchEGraph building-energy dataset for graph-based and weather-conditioned learning.
Dataset Summary
Total cases in manifest.csv: 300
Unique buildings: 75
Unique weather IDs: 48
n_steps: always 8,760
n_spaces range: 2 to 132
This package currently stores:
manifest.csv (index of all demo cases)
building/ (75 files)
geometry/ (75 files)
weather/ (48 files)
energy/ (300 files)
split/ (demo split CSV files)… See the full description on the dataset page: https://huggingface.co/datasets/ArchEGraph/ArchEGraph-demo.arched-halls-decision-matrix
Arched Halls Decision Matrix / Macierz decyzyjna hal łukowych
Dataset summary
This Polish-language dataset describes 30 practical application scenarios for an arched hall (hala łukowa) in agriculture, storage, logistics, transport, industry, waste management, infrastructure, sports, public facilities, seasonal buildings, construction and energy.
Each record connects the intended use of an arched hall with qualitative decision factors such as indoor climate… See the full description on the dataset page: https://huggingface.co/datasets/halalukowa24/arched-halls-decision-matrix.industrial-technical-archive
🚀 Latest Updates (July, 2026)
Version: v07.2026 (Verified)
Status: Integrated with 1,000,000+ records.
New Files: product-E-26-07-2026.csv & product-V-26-07-2026.csv.
QTE Technologies: Industrial & Scientific Knowledge Base
Wikidata Entity: Q138411149
IPFS CID: bafybeibogxxuhmzfrsuhcfd4qr4tmc4okhmrcwhp3266hq47ccuyjnjxoq
Official Neural Hub: qtetech.github.io
This is the permanent technical archive for QTE Technologies, ensuring long-term accessibility of… See the full description on the dataset page: https://huggingface.co/datasets/QTE-Technologies/industrial-technical-archive.sinhala-tts-dataset-archive-20260429-082457
Sinhala TTS Dataset
Clean, segmented Sinhala speech from the "Unlimited History" YouTube series by @sunchare.
Stats
Metric
Value
Utterances
218
Train
208
Val
10
Hours
0.51
Mean duration
8.5s
Sample rate
22050 Hz
Pipeline
Raw YouTube audio -> HTDemucs -> VoiceFixer + DeepFilterNet3 ->
Diarization -> Silero-VAD -> ASR (faster-whisper: C:\Users\kosal\sinhala-tts\whisper-small-si-ct2) -> Quality filtering (SNR>=20.0dB)
Format… See the full description on the dataset page: https://huggingface.co/datasets/outlawmold/sinhala-tts-dataset-archive-20260429-082457.cars_from_drom.ru_archive_2007-2025More information on the parsing process can be found here: https://github.com/zavzyatiy/drom_archive_parser.
This dataset is also published on Kaggle: https://www.kaggle.com/datasets/assaabramovich/resaled-cars-from-drom-ruarchive-2018-2023/.
Main dataset with all data: drom_archive_2007-2025_full.csv
Dataset with (almost) all configurations from Drom for cars in data: additional_data/drom-24-07-2025-all_main_cars_configurations.csv
Dataset with identification of regions for all cities in… See the full description on the dataset page: https://huggingface.co/datasets/zavzyatiy/cars_from_drom.ru_archive_2007-2025.ARC-Challenge-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in ARC Challenge. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
prog-archivesjazz-music-archivestop-1M-websitearc_agi_2_human_testing
ARC-AGI-2 Human testing data
This file contains data from human testing sessions on ARC-AGI tasks.
Each row represents a single test attempt by a human participant on a specific task-test pair in the "Public Train" or "Public Eval" ARC-AGI-2 datasets. Not all tasks in the released "Public Train"
sets were tested, so these results are not comprehensive. This data does not include tasks from "Semi Private Evaluation" or "Private Evaluation"
Column Descriptions… See the full description on the dataset page: https://huggingface.co/datasets/arcprize/arc_agi_2_human_testing.ARC-Easy-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in ARC-Easy. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
ArcMMLU
Introduction
ArcMMLU is a Chinese benchmark specifically designed for evaluating LLMs on Library & Information Science (LIS). It aims to evaluate the knowledge and reasoning capabilities of LLMs in the LIS academic field, which covers four key sub-areas: Archival Science, Data Science, Library Science, and Information Science. Please refer to our paper for more information ArcMMLU: A Library and Information Science Benchmark for Large Language Models
It is important to note that the… See the full description on the dataset page: https://huggingface.co/datasets/patrickshitou/ArcMMLU.arc-agi-3-schema-traces-gpt56
ARC-AGI-3 Schema Gameplay Trajectories — GPT-5.6 Sol
This release contains every gpt-5.6-sol gameplay trajectory produced on our
cluster with the world_model_v5 agent harness — 100 runs across the 25 public
ARC-AGI-3 games — plus a dependency-free scoring utility.
It is the GPT-5.6 Sol member of a family built by the same harness and the same
sanitizer, so trajectories can be compared game by game:
arc-agi-3-schema-traces-fable5 — Claude Fable 5, best per game (25)… See the full description on the dataset page: https://huggingface.co/datasets/guanning/arc-agi-3-schema-traces-gpt56.bbench-dep-song-describer
The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation
Ilaria Manco*1,2,
Benno Weck*3,
Seungheon Doh4,
Minz Won5,
Yixiao Zhang1,
Dmitry Bogdanov3,
Yusong Wu6,
Ke Chen7,
Philip Tovstogan3,
Emmanouil Benetos1,
Elio Quinton2,
George Fazekas1,
Juhan Nam4
1 QMUL, 2 UMG, 3 UPF, 4 KAIST, 5 ByteDance, 6 MILA, 7 UCSD
* equal contribution
This repository contains starter code for the Song Describer Dataset (SDD).
Paper (accepted to the ML… See the full description on the dataset page: https://huggingface.co/datasets/Archit00/bbench-dep-song-describer.arcs-authority-vulnerability
ARCS Authority Vulnerability Evaluation Dataset v1.1
Description
Empirical evaluation data measuring authority vulnerability in AI systems. Covers single-model evaluation, two-hop agent chain propagation, and three-hop agent chain propagation across six independent AI lineages.
This is the first published dataset measuring:
Whether AI models accept false authority claims under adversarial pressure
Whether authority vulnerability propagates between models in… See the full description on the dataset page: https://huggingface.co/datasets/aa8899/arcs-authority-vulnerability.polaris-arctic-v1
Arctic v1
Arctic is one of six datasets in Polaris: Learning to Generate Table Descriptions from Retrieval
Feedback, alongside aw, lter, ecir, wikitables, and
wtr.
It holds 251 tables sampled at random from the Environmental Data Initiative (EDI), a repository of
long-term ecological research data — lake water temperature, soil chemistry, rainfall, coral
taxonomy — and 20 keyword queries over them. For each query–table pair, a person decided whether that
table answers that… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-arctic-v1.polaris-arctic-v2
Arctic v2
Arctic is one of six datasets in Polaris: Learning to Generate Table Descriptions from Retrieval
Feedback, alongside aw, lter, ecir, wikitables, and
wtr.
It holds 251 tables sampled at random from the Environmental Data Initiative (EDI), a repository of
long-term ecological research data — lake water temperature, soil chemistry, rainfall, coral
taxonomy — and 20 keyword queries over them. For each query–table pair, a person decided whether
that
table answers that… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-arctic-v2.llm-cold-start-benchmark
LLM Container Cold-Start Benchmark
Measurements of how long it takes to bring a language model from cold storage to
a state where it can serve its first token, across 25 open-weight
models spanning 17 architecture families and
100.9 GiB of checkpoints, on a single NVIDIA T4.
Cold start is the latency a serverless or scale-to-zero inference platform pays
when it has no warm replica. It decomposes into weight transfer from storage,
deserialization into host memory, transfer to the… See the full description on the dataset page: https://huggingface.co/datasets/ArchCoder/llm-cold-start-benchmark.data_jobs
🧠 data_jobs Dataset
A dataset of real-world data analytics job postings from 2023, collected and processed by Luke Barousse.
Background
I've been collecting data on data job postings since 2022. I've been using a bot to scrape the data from Google, which come from a variety of sources.
You can find the full dataset at my app datanerd.tech.
Serpapi has kindly supported my work by providing me access to their API. Tell them I sent you and get 20% off paid plans.… See the full description on the dataset page: https://huggingface.co/datasets/archanaj/data_jobs.
