datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai2_arc
Dataset Card for "ai2_arc"
Dataset Summary
A new dataset of 7,787 genuine grade-school level, multiple-choice science questions, assembled to encourage research in
advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains
only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. We are also
including a corpus of over 14 million science sentences… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ai2_arc.jfk-archives
Dataset Card for JFK Archives
This dataset is a collection of all records pertaining to the assassination of the
US president, John F. Kennedy, released until April 2025 through archives.org
by the US government.
Dataset Details
Dataset Description
The original data downloaded from archives.org
consists of 56,300 scanned documents in PDF format, released until April 2025. The files are
organized by their release year(s): 2107-2018, 2021, 2022, 2023 and 2025.… See the full description on the dataset page: https://huggingface.co/datasets/farhanhubble/jfk-archives.m_arc
Multilingual ARC
Dataset Summary
This dataset is a machine translated version of the ARC dataset.
The Icelandic (is) part was translated with Miðeind's Greynir model and Norwegian (nb) was translated with DeepL. The rest of the languages was translated using GPT-3.5-turbo by the University of Oregon, and this part of the dataset was originally uploaded to this Github repository.
tinyAI2_arc
tinyAI2_arc
Welcome to tinyAI2_arc! This dataset serves as a concise version of the AI2_arc challenge dataset, offering a subset of 100 data points selected from the original compilation.
tinyAI2_arc is designed to enable users to efficiently estimate the performance of a large language model (LLM) with reduced dataset size, saving computational resources
while maintaining the essence of the ARC challenge evaluation.
Features
Compact Dataset: With only 100 data… See the full description on the dataset page: https://huggingface.co/datasets/tinyBenchmarks/tinyAI2_arc.msmarco-v2.1-snowflake-arctic-embed-l
Snowflake Arctic Embed L Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using Snowflake's Arctic Embed L and are intended to serve as a simple baseline for dense retrieval-based methods.
Retrieval Performance
Retrieval performance for the TREC DL21-23, MSMARCOV2-Dev and Raggy Queries can be found below with BM25 as a baseline. For both… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-l.omnimcp_python_backend_architect_teaser
🚀 OmniMCP: Python Backend Architect (Evaluation Teaser Edition)
🛡️ DON'T WANT TO TRAIN RAW DATASETS? RUN IT IN CURSOR & CLAUDE TODAY!
You don't need A100 GPUs, Axolotl, or complex Unsloth fine-tuning. This package now includes a Turnkey Ready-to-Run Model Context Protocol (MCP) Server that plugs directly into Cursor IDE and Claude Desktop with 1-click!
⚡ What the Turnkey MCP Guard does inside Cursor & Claude in <100ms:
🩺 Instant Traceback Diagnosis: Feed any… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_python_backend_architect_teaser.msmarco-v2.1-snowflake-arctic-embed-m-v1.5
Snowflake Arctic Embed M V1.5 Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using Snowflake's Arctic Embed M v1.5 and are intended to serve as a simple baseline for dense retrieval-based methods.
It's worth noting that Snowflake's Arctic Embed M v1.5 is optimized for efficient embeddings and thus supports embedding truncation and quantization. More… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-m-v1.5.arc-cot
Augmented ARC-Challenge Dataset with Chain-of-Thought Reasoning
Dataset Description
This dataset was created by augmenting the train subset of the AI2 Reasoning Challenge (ARC) dataset with chain-of-thought reasoning generated by Google's Gemini Pro language model. The goal is to provide additional context and intermediate reasoning steps to help models better solve the challenging multiple-choice science questions in ARC.
Dataset Structure
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/Locutusque/arc-cot.USCode-QAPairs-Finetuning
USCode-QueryPairs Dataset
This dataset contains query-answer pairs curated from the United States Code, suitable for fine-tuning any embedding model. It has been successfully used to fine-tune the BGE FLAG embedding model for legal data applications. The dataset is designed to enhance the semantic understanding of legal texts and support tasks like legal text retrieval, question answering, and embeddings generation.
Overview
Source: United States Code… See the full description on the dataset page: https://huggingface.co/datasets/ArchitRastogi/USCode-QAPairs-Finetuning.Nemotron-RL-ARC-AGI-v1
Dataset Description:
Nemotron-RL-ARC-AGI-v1 is a reinforcement-learning (RL) gym environment dataset of single-step ARC-AGI puzzle prompts intended for RL post-training of large language models. Each row corresponds to one ARC puzzle (a set of (input grid, output grid) demonstration pairs plus a single test input grid) rendered as a text prompt; reward is binary (1.0 / 0.0) determined by exact-match comparison against the ground-truth output grid. No LLM judge is used, no… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-ARC-AGI-v1.arc_de
ARC (DE) — Boldt German Evaluation Suite
Improved German translations of the ARC-Easy and ARC-Challenge benchmarks (Clark et al., 2018), part of the Boldt German Evaluation Suite.
Translation
The validation splits of the datasets were translated from the English originals using Tower+ 72B. Rather than translating question and answer options separately, we translated complete question-answer instances end-to-end to preserve internal consistency and task logic. A small… See the full description on the dataset page: https://huggingface.co/datasets/Boldt/arc_de.arcd
Dataset Card for "arcd"
Dataset Summary
Arabic Reading Comprehension Dataset (ARCD) composed of 1,395 questions posed by crowdworkers on Wikipedia articles.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
plain_text
Size of downloaded dataset files: 1.94 MB
Size of the generated dataset: 1.70 MB
Total amount of disk used: 3.64 MB
An… See the full description on the dataset page: https://huggingface.co/datasets/hsseinmz/arcd.razavi-benchRazavi-bench
An expert-curated benchmark for analog-design reasoning.
Razavi-bench packages the question-answer assessments from Behzad Razavi's
Analog Design Experiments With AI Part 1 and Part 2 into a clean
one-task-per-directory benchmark. The tasks probe whether a model can reason
about MOS devices, small-signal circuits, feedback, oscillators, comparators,
dividers, LNAs, TIAs, and LC oscillators.
Each task directory keeps only the benchmark prompt, figure, and curated… See the full description on the dataset page: https://huggingface.co/datasets/Arcadia-2026/razavi-bench.MathNet
Quick Start · Overview · Tasks · Comparison · Dataset Stats · Data Sources · Pipeline · Schema · License · Citation
This is the official MathNet v0. A larger version v1 will be uploaded soon (more countires, problems and richer metadata). Schema is stable but field values may be revised in v1.
Quick start
from datasets import load_dataset
# Default: all problems
ds = load_dataset("ShadenA/MathNet", split="train")
# Or a specific country / competition-body config
arg… See the full description on the dataset page: https://huggingface.co/datasets/archya/MathNet.securecode-web-archive
SecureCode Web: Traditional Web & Application Security Dataset
Production-grade web security vulnerability dataset with complete incident grounding, 4-turn conversational structure, and comprehensive operational guidance
Paper | GitHub | Dataset | Model Collection | Blog Post
What's new in v2.6
v2.6 restores proper Express.js coverage for the topics whose examples were removed in v2.5.1 (they had
shared one reused answer). 29 new, genuinely distinct Express.js… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/securecode-web-archive.uhura-arc-easy
Dataset Card for Uhura-Arc-Easy
Dataset Summary
Uhura-ARC-Easy is a widely recognized scientific question answering benchmark composed of multiple-choice science questions derived from grade-school examinations that test various styles of knowledge and reasoning.
The original English version of the benchmark originates from Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge (Clark et al., 2018) and is divided into "Challenge" and "Easy"… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/uhura-arc-easy.ARCHIVE-TEXT-URLS
Internet Archive English Text URLs Dataset
Dataset Description
This dataset contains 11,151,637 direct download URLs to OCR-processed text files from the Internet Archive's digital library. All entries are English-language texts spanning books, documents, historical records, and various other written materials.
Dataset Summary
Total Rows: 11,151,637
Language: English
Source: Internet Archive
Format: CSV with metadata and direct text file URLs
Text… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/ARCHIVE-TEXT-URLS.ai2_arc-pt
marlosb/ai2-arc-pt
This dataset is a Portuguese translation of the original AI2 ARC (AI2 Reasoning Challenge) dataset released by AllenAI.
Original Dataset
Hugging Face: allenai/ai2_arc
Homepage: https://allenai.org/data/arc
Paper: Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Dataset Summary
AI2 ARC consists of 7,787 genuine grade-school level, multiple-choice science questions, assembled to encourage research in advanced… See the full description on the dataset page: https://huggingface.co/datasets/marlosb/ai2_arc-pt.policystrategies-archive
🏛️ Open-Source Macro-Strategy, Financial History & Intelligence Archive
🌐 Overview & Institutional Mission
This public repository serves as the official open-source knowledge graph and metadata registry for r/policystrategies.
We aggregate, document, and cross-reference declassified historical intelligence dossiers, sovereign debt crises, systemic market manipulations, and geoeconomic conflicts using verified open-source intelligence (OSINT) and primary… See the full description on the dataset page: https://huggingface.co/datasets/stratigahq/policystrategies-archive.hotpot_qa_archiveHotpotQA is a new dataset with 113k Wikipedia-based question-answer pairs with four key features:
(1) the questions require finding and reasoning over multiple supporting documents to answer;
(2) the questions are diverse and not constrained to any pre-existing knowledge bases or knowledge schemas;
(3) we provide sentence-level supporting facts required for reasoning, allowingQA systems to reason with strong supervisionand explain the predictions;
(4) we offer a new type of factoid comparison questions to testQA systems’ ability to extract relevant facts and perform necessary comparison.cbi-archive-corpus
Central Bank of Ireland Public Archive Corpus
A page-anchored, provenance-classified corpus of the Central Bank of Ireland's
public document archive. 5,568 documents and 89,242 page or pseudo-page rows.
PDF rows have true source-page anchors; most Office and archive rows do not.
This is an unofficial derived work. It is not published by, affiliated with, or
endorsed by the Central Bank of Ireland.
What makes this different from a pile of scraped PDFs
Two things.… See the full description on the dataset page: https://huggingface.co/datasets/aditya487/cbi-archive-corpus.ARC-eu
Dataset Card for ARC-eu
Point of Contact: hitz@ehu.eus
Dataset Description
Dataset Summary
ARC-eu is the professional translation to Basque of ARC's
(Clark et al., 2018) validation and test partitions.
ARC is a QA benchmark of grade-school level, multiple-choice science questions.
Languages
eu-ES
Dataset Structure
Data Instances
ARC-eu examples look like this:
{
"id": "MCAS_2000_4_6",
"question": "Zein teknologia… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/ARC-eu.Word-Puzzles-ARC-Unique-50000
Word-Puzzles-ARC-Unique-46000
This dataset is a synthetic 46,000-row word-puzzle corpus focused on answerable reasoning tasks with explicit gold answers.
Version
This upload corresponds to the harder v2 build.
Bucket mix
15,000 formal deduction
12,500 constraint-based lexical deduction
10,000 symbolic substitution
7,500 semantic association
1,000 riddles
Hardening changes in v2
formal puzzles use 6 entities instead of 5
lexical… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Word-Puzzles-ARC-Unique-50000.Arc-ATLAS-Teach-v1
Arc-ATLAS-Teach
Summary
This revision bundles 624 high-quality adaptive teaching examples that were generated and validated with the latest five-pass pipeline. Every dialogue walks through the full instructional arc—probe, draft plan, checkpoint feedback, revised plan, and final solution—so the teaching policy observes the complete adjustment process without ever seeing the canonical answer. Probe turns capture the student’s diagnostic attempt, teacher plans and… See the full description on the dataset page: https://huggingface.co/datasets/Arc-Intelligence/Arc-ATLAS-Teach-v1.arc-trThis Dataset is part of a series of datasets aimed at advancing Turkish LLM Developments by establishing rigid Turkish benchmarks to evaluate the performance of LLM's Produced in the Turkish Language.
Dataset Card for arc-tr
malhajar/arc-tr is a translated version of arc aimed specifically to be used in the OpenLLMTurkishLeaderboard
This Dataset contains rigid tests extracted from the paper Think you have Solved Question Answering?
Developed by: Mohamad Alhajar
Data… See the full description on the dataset page: https://huggingface.co/datasets/malhajar/arc-tr.arc_ca
Dataset Card for arc_ca
arc_ca is a question answering dataset in Catalan, professionally translated from the Easy and Challenge versions of the ARC dataset in English.
Dataset Details
Dataset Description
arc_ca (AI2 Reasoning Challenge - Catalan) is based on multiple-choice science questions at elementary school level. The dataset consists of 2950 instances in the Easy version (570 in the test and 2380 instances in the validation split) and… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/arc_ca.ARC_AGI_V1_ULTRASNU_Ko-ARC
Note: Evaluation code and task configurations for this benchmark are available in the evaluation_code directory of the Korean Benchmark Suite GitHub repository. The evaluation setup is built on the Language Model Evaluation Harness and supports standardized model assessment.
Dataset Card for Ko-ARC
Dataset Summary
Ko-ARC is a Korean adaptation of the AI2 Reasoning Challenge (ARC) dataset.
It consists of multiple-choice science questions designed to assess… See the full description on the dataset page: https://huggingface.co/datasets/thunder-research-group/SNU_Ko-ARC.ARC-poly
ARC Polyglot
This dataset is a multilingual version of the original ARC (AI2 Reasoning Challenge) dataset, which consists of multiple-choice questions designed to test the reasoning abilities of AI systems. The polyglot version includes translations of the original English questions into various languages, allowing for evaluation of language models across different linguistic contexts.
All languages supported
* ar (Arabic)
* bn (Bengali)
* ca (Catalan)
* da (Danish)
* de (German)… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/ARC-poly.ARC-Challenge-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in ARC Challenge. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
