datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AA-LCR
Artificial Analysis Long Context Reasoning (AA-LCR) Dataset
AA-LCR includes 100 hard text-based questions that require reasoning across multiple real-world documents, with each document set averaging ~100k input tokens. Questions are designed such that answers cannot be directly retrieved from documents and must instead be reasoned from multiple information sources.
New in Version 1.1 (September 2026)
Sixteen corrected answer keys. Each one was re-verified… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-LCR.AA-Omniscience-Public
Public Dataset for AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models
AA-Omniscience-Public contains 600 questions across a wide range of domains used to test a model’s knowledge and hallucination tendencies.
Leaderboard and detailed results
Paper
Introduction
We introduce AA-Omniscience, a benchmark dataset designed to measure a model’s ability to both recall factual information accurately across domains, and correctly… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-Omniscience-Public.spreadsheet-arena-release
Spreadsheet Arena
A dataset of 555 pairwise human preference votes over LLM-generated spreadsheets, spanning 124 distinct user-submitted prompts and 17 models.
This is the public release accompanying the Spreadsheet Arena paper.
Contents
battles.csv
models.csv
outputs/<id>/
sheet.json
sheet.xlsx
<id> is a 16-char hex identifier (HMAC-SHA256 of an internal UUID under a… See the full description on the dataset page: https://huggingface.co/datasets/Longitude-Labs/spreadsheet-arena-release.ML-ArXiv-PapersThis dataset contains the subset of ArXiv papers with the "cs.LG" tag to indicate the paper is about Machine Learning.
The core dataset is filtered from the full ArXiv dataset hosted on Kaggle: https://www.kaggle.com/datasets/Cornell-University/arxiv. The original dataset contains roughly 2 million papers. This dataset contains roughly 100,000 papers following the category filtering.
The dataset is maintained by with requests to the ArXiv API.
The current iteration of the dataset only contains… See the full description on the dataset page: https://huggingface.co/datasets/CShorten/ML-ArXiv-Papers.ArabicMMLU
Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin
MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne
Introduction
We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/ArabicMMLU.arabic-stem-lexicon
Arabic Diacritized-Stem Lexicon
An undiacritized Arabic surface form → its most frequent diacritized stem.
Standard Arabic writes no short vowels, so anything that has to pronounce Arabic
must first put them back. A neural diacritizer does that well on rare words, where
inference is the only thing there is. On common words it is the wrong tool:
which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked
up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.arena-human-preference-55kDataset for Kaggle competition on predicting human preference on Chatbot Arena battles.
The training dataset includes over 55,000 real-world user and LLM conversations and user preferences across over 70 state-of-the-art LLMs, such as GPT-4, Claude 2, Llama 2, Gemini, and Mistral models.
Each sample represents a battle consisting of 2 LLMs which answer the same question, with a user label of either prefer model A, prefer model B, tie, or tie (both bad).
Citation
Please cite the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-human-preference-55k.sr-artifact-prominence
SR Artifact Prominence
Annotated super-resolution artifact regions across four image subsets, with
crowdsourced per-region prominence scores, artifact type labels, and
natural-language descriptions.
Prominence is the fraction of valid crowd workers who answered that the
highlighted region contains a noticeable super-resolution artifact.
Subsets
Subset
Source dataset
Source images
Masks
Notes
open_images
Open Images
547
1,523
GT + LR-bicubic + multiple SR… See the full description on the dataset page: https://huggingface.co/datasets/imolodetskikh/sr-artifact-prominence.argument_quality_ranking_30k
Dataset Card for Argument-Quality-Ranking-30k Dataset
Dataset Summary
Argument Quality Ranking
The dataset contains 30,497 crowd-sourced arguments for 71 debatable topics labeled for quality and stance, split into train, validation and test sets.
The dataset was originally published as part of our paper: A Large-scale Dataset for Argument Quality Ranking: Construction and Analysis.
Argument Topic
This subset contains 9,487 of the arguments only with… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/argument_quality_ranking_30k.TTS_ArenaTTS Arena's DB is SQLlite DB file. The above is just a summary query that should be useful for TTS developers to evaluate faults of their model.
Why no audio samples?
Unsafe. Cannot constantly oversee the output of uncontrolled HuggingFace Spaces. While it could be safeguarded by using an ASR model before uploading, something unwanted may still slip through.
Useful queries for TTS developers and evaluators
All votes mentioning specified TTS model:… See the full description on the dataset page: https://huggingface.co/datasets/Pendrokar/TTS_Arena.arc-agi-3-schema-traces
ARC-AGI-3 Schema Gameplay Trajectories
This release contains 50 ARC-AGI-3 gameplay trajectories and a dependency-free
scoring utility. The trajectories are split evenly across two collections:
gpt_5_6_sol/: 25 GPT-5.6 Sol trajectories.
claude_fable_opus/: 25 trajectories from Claude Opus 4.8 and Claude Fable 5.
Each trajectory directory includes run.json, a streamed events.jsonl event
log, sanitized session data, snapshots, and the shareable text/image files
produced during… See the full description on the dataset page: https://huggingface.co/datasets/schema-harness/arc-agi-3-schema-traces.arguana-qrels
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/arguana-qrels.arct
The Argument Reasoning Comprehension Task: Identification and Reconstruction of Implicit Warrants
https://github.com/UKPLab/argument-reasoning-comprehension-task
@InProceedings{Habernal.et.al.2018.NAACL.ARCT,
title = {The Argument Reasoning Comprehension Task: Identification
and Reconstruction of Implicit Warrants},
author = {Habernal, Ivan and Wachsmuth, Henning and
Gurevych, Iryna and Stein, Benno},
publisher = {Association for… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/arct.LLM-Artifacts
Under the Surface: Tracking the Artifactuality of LLM-Generated Data
Debarati Das†¶, Karin de Langis¶, Anna Martin-Boyle¶, Jaehyung Kim¶, Minhwa Lee¶, Zae Myung Kim¶
Shirley Anugrah Hayati, Risako Owan, Bin Hu, Ritik Sachin Parkar, Ryan Koo,
Jong Inn Park, Aahan Tyagi, Libby Ferland, Sanjali Roy, Vincent Liu
Dongyeop Kang
Minnesota NLP, University of Minnesota Twin Cities
† Project Lead,
¶ Core Contribution,
Arxiv
Project Page
📌 Table of Contents
Introduction… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/LLM-Artifacts.cbi-archive-raw
Central Bank of Ireland Archive: original source files
6,309 original files, 6.56 GB. Every PDF, spreadsheet, Word document and
archive gathered from the Central Bank of Ireland's public website, stored by
content hash so that a search result can be turned back into the document a
human would actually read.
This is the raw tier. If you want the text, you almost certainly want
aditya487/cbi-archive-corpus
instead: 5,568 documents and 89,242 page or pseudo-page rows as Parquet… See the full description on the dataset page: https://huggingface.co/datasets/aditya487/cbi-archive-raw.w2t-llm-arc-easy-lora
W2T Llm Arc Easy Lora
This repository contains artifacts for the W2T paper:
Paper: W2T: LoRA Weights Already Know What They Can Do
Repo: Weight2Token
Summary
ARC-Easy LoRA checkpoints and prepared metadata used for performance prediction.
Source Status
Storage location: local
Verification status: confirmed
Files
See manifest.json for the exact local or remote source paths used to prepare this release.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/Xiaolong-Han/w2t-llm-arc-easy-lora.cryptocurrency-futures-ohlcv-dataset-1marxiv_nlp
arXiv Abstracts
Abstracts for the cs.CL category of ArXiv between 1991 and 2024. This dataset was created as an instructional tool for the Clustering and Topic Modeling chapter in the upcoming
"Hands-On Large Language Models" book.
The original dataset was retrieved here.
This subset will be updated towards the release of the book to make sure it captures relatively recent articles in the domain.
CADS-dataset
CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography
Overview
CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems.
The framework consists of two main components:
CADS-dataset:
22,022 CT volumes with complete annotations for 167 anatomical structures.
Most extensive whole-body CT dataset… See the full description on the dataset page: https://huggingface.co/datasets/arekborucki/CADS-dataset.chatbot-arena-elo
LMSYS Chatbot Arena ELO Scores
This dataset is a datasets-friendly version of Chatbot Arena ELO scores,
updated daily from the leaderboard API at
https://huggingface.co/spaces/lmarena-ai/chatbot-arena-leaderboard.
Updated: 20250717
Loading Data
from datasets import load_dataset
dataset = load_dataset("mathewhe/chatbot-arena-elo", split="train")
The main branch of this dataset will always be updated to the latest ELO and
leaderboard version. If you need a fixed dataset… See the full description on the dataset page: https://huggingface.co/datasets/mathewhe/chatbot-arena-elo.eu-ai-act-article-50-scoreboard
Article 50 historical public-evidence snapshot
This work was produced through an AI-assisted workflow directed by the author. Historical work used Anthropic assistance; the retrospective correction uses OpenAI GPT-6, with separate bounded Gemini advice. All three providers have products in the scored set.
Purpose: provide the corrected paper's version 1.1 bundle under v1_1. Start with its README and correction note. The paper and deposit and GitHub repository identify the same… See the full description on the dataset page: https://huggingface.co/datasets/NMAIResearch/eu-ai-act-article-50-scoreboard.us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites
Historical US layoffs archive: 6,799 WARN Act notices that state websites no longer list (2000-2025), recovered
Rebuilt 2026-09-25. Five state labor agencies — Connecticut, Michigan, New York,
North Carolina and Pennsylvania — retired the web pages their older WARN Act
layoff notices lived on. Their current pages start years later. This dataset is
every notice in our file that came from one of those retired pages and is not
on the agency's live page today: 6,799 notices, 6,799… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites.ArFake
ArFake-Dataset
ARFAKE: A Robust Framework for Multi-Dialect Arabic Speech Spoofing Detection Benchmark
ARFAKE is the first end-to-end benchmark for Arabic speech spoofing detection across multiple dialects. The framework systematically generates synthetic Arabic speech, evaluates its intelligibility and realism, constructs a large-scale spoofing dataset, trains robust detectors, and evaluates generalization across both unseen generators and unseen dialects.… See the full description on the dataset page: https://huggingface.co/datasets/Mohammed01/ArFake.arabic-dialects-gold20
arabic-dialects-gold20
660 sentences: 33 Arabic lects × 20 sentences, each with fully diacritized
dialectal orthography, the undiacritized surface form, gold IPA, an engine
draft, an English gloss, machine-verified phonetic feature tags, per-row
verification metadata, and notes citing the dialectological literature that
grounds the row.
Columns (TSV, UTF-8, one file per lect):
id, sentence, raw, ipa, ipa_o2i, gloss_en, features, notes, fable_corrections, verification… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-dialects-gold20.us-layoffs-by-metro-area-msa-warn-act
US layoffs by metro area: 54,225 WARN notices mapped to 765 metro and micro areas
Rebuilt 2026-09-24. 765 of the 935 US core-based statistical areas carry at least one
layoff notice on record — 361 metropolitan and 404 micropolitan.
Nobody hires, sells or reports by county. A recruiter covers Austin; an account team books the
Phoenix metro; a reporter writes Bay Area layoffs. State agencies publish neither — they publish
the site of a layoff as free text in 48 different… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-layoffs-by-metro-area-msa-warn-act.ScienceQAThis is the ScientificQA dataset by Saikh et al (2022).
@article{10.1007/s00799-022-00329-y,
author = {Saikh, Tanik and Ghosal, Tirthankar and Mittal, Amish and Ekbal, Asif and Bhattacharyya, Pushpak},
title = {ScienceQA: A Novel Resource for Question Answering on Scholarly Articles},
year = {2022},
journal = {Int. J. Digit. Libr.},
month = {sep}
}
artofproblemsolvingopenvino-arc140v-lunarlake
OpenVINO local-inference on an Intel Arc 140V (Lunar Lake) iGPU
Reference performance data for running local models on a single Intel Core Ultra 7 258V
(Lunar Lake) laptop with the integrated Intel Arc 140V (Xe2) GPU, via OpenVINO. All
inference runs on the iGPU; the NPU stays idle throughout, confirmed by the telemetry here.
This is reference characterization shared by a non-expert contributor — careful measurements
on one machine, offered so others can compare and correct, not… See the full description on the dataset page: https://huggingface.co/datasets/blairducrayoppat/openvino-arc140v-lunarlake.gpqa_diamondArcBench
ArcBench: ML Conference Oral Paper-Presentation Benchmark
This benchmark is from the paper Narrative-Driven Paper-to-Slide Generation via ArcDeck.
A curated benchmark dataset of 100 oral presentation paper-slide deck link pairs from top-tier machine learning conferences (CVPR, ICCV, ICLR, ICML, NeurIPS), spanning 2022–2025. Each entry provides rich metadata together with links to the original paper PDF and presentation slides, plus a script that downloads them all in one step.… See the full description on the dataset page: https://huggingface.co/datasets/ArcDeck/ArcBench.
