datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sql-create-context
Overview
This dataset builds from WikiSQL and Spider.
There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/sql-create-context.context_qa_sum_qwen3_synthetic
Context-based QA and Summarization Synthetic Dataset
Overview
This dataset contains synthetic context-based question-answering (QA) and summarization data. The data was synthesized using:
Source context: openbmb/Ultra-FineWeb
Synthesis model: Qwen3-30B-A3B-Instruct-2507
Each context is obtained by taking the initial segment of raw pretraining text from Ultra-FineWeb, truncated to at most the corresponding number of tokens, while ensuring the truncation does not occur in… See the full description on the dataset page: https://huggingface.co/datasets/yuyijiong/context_qa_sum_qwen3_synthetic.ledger-long-context-KPI-QA
LEDGER — Long-Context KPI Question Answering & Page Retrieval
This dataset is part of the LEDGER (Long-context Evaluation of Documents for
Grounded Extraction and Retrieval) benchmark.
It supports two of the three LEDGER tasks:
Page-level KPI retrieval — given a natural-language question about a financial
KPI and the corresponding annual report, retrieve the relevant page(s). Each row
includes TREC-style graded relevance judgments (qrels) over all candidate pages.… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-KPI-QA.persona-drift-contextecho
ContextEcho — Released Dataset
Per-cell evaluation corpus and donated session prefixes for the ContextEcho
benchmark. This Hugging Face repository hosts the released dataset artifacts.
The canonical project page, latest README, code, reproduction instructions, and
donation workflow are maintained on GitHub:
https://github.com/Accenture/ContextEcho
Donate a coding-agent session: https://accenture.github.io/ContextEcho/donate/
For the formal datasheet, see DATASHEET.md.… See the full description on the dataset page: https://huggingface.co/datasets/contextecho2026/persona-drift-contextecho.Multi-turn_Long-context_Benchmark_for_LLMs
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
Arxiv: https://www.arxiv.org/abs/2507.13681
Huggingface: https://huggingface.co/papers/2507.13681
Introduction
LoopServe Multi-Turn Dialogue Benchmark is a comprehensive evaluation dataset comprising multiple diverse datasets designed to assess large language model performance in realistic conversational scenarios.
Unlike traditional benchmarks that place queries only at the end… See the full description on the dataset page: https://huggingface.co/datasets/TreeAILab/Multi-turn_Long-context_Benchmark_for_LLMs.mix-context-post-training-128k
Mix-Context Post-Training Dataset for 128K Context Extension
Overview
Mix-Context Post-Training 128K is a dataset designed specifically for post-training context window extension of pretrained LLMs.
It targets the stage after base pretraining, where a model is adapted to operate over much longer contexts (up to 128K tokens) while preserving short-context behavior. The dataset mixes short- and long-context packed sequences with a controlled length distribution to support:… See the full description on the dataset page: https://huggingface.co/datasets/ghostcc3/mix-context-post-training-128k.ContextualIntegritySyntheticDataset
Contextual Integrity Synthetic Dataset
This repository contains the synthetic dataset introduced in the paper "Contextual Integrity in LLMs via Reasoning and Reinforcement Learning".
Paper | Code | Blog
Dataset Summary
The Contextual Integrity (CI) synthetic dataset consists of 729 examples featuring diverse contexts and information disclosure norms. It is designed to instill reasoning capabilities in LLMs regarding what information is appropriate to share while… See the full description on the dataset page: https://huggingface.co/datasets/huseyinatahaninan/ContextualIntegritySyntheticDataset.austen-corpus
ContextLab Jane Austen Corpus
Dataset Description
This dataset contains works of Jane Austen (1775-1817), preprocessed for computational stylometry research. The texts were sourced from Project Gutenberg and cleaned for use in the paper "A Stylometric Application of Large Language Models" (Stropkay et al., 2025).
The corpus includes 7 books by Jane Austen, including Pride and Prejudice, Sense and Sensibility, and Emma. All text has been converted to lowercase and cleaned… See the full description on the dataset page: https://huggingface.co/datasets/contextlab/austen-corpus.ContextEval
Contextualized Evaluations: Taking the Guesswork Out of Language Model Evaluations
Dataset Summary
We provide here the data accompanying the paper: Contextualized Evaluations: Taking the Guesswork Out of Language Model Evaluations.
Dataset Structure
Data Instances
We release the set of queries, as well as the autorater & human evaluation judgements collected for our experiments.
Data overview
List of queries: Data Structure
The list… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ContextEval.long_context_hindi
Dataset
This dataset was filtered from AI4BHarat dataset sangraha,which is the largest high-quality, cleaned Indic language pretraining data containing 251B tokens summed up over 22 languages, extracted from curated sources, existing multilingual corpora and large scale translations.
This dataset contains only Hindi as of now
Information
First this dataset is mainly for long context training
The minimum len is 6000 and maximum len is 3754718
Getting started… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/long_context_hindi.Synthetic-Dataset-Childrens-Stories**Status: released 13-09-2026, repacked 14-09-2026.** The 14-09-2026 repack replaced 58
items after the acceptance gates were strengthened (prompt-instruction leaks, markdown
bullet lists and blockquotes); the other 29,942 are unchanged. Development stopped, pipeline released 17/09/26.
SAMPLE RELEASE: 30,000 synthetic children's short stories for early-reader language modelling.
Metrics
Value
genres
26
stories per genre
1.153-1.154K stories
total characters
38… See the full description on the dataset page: https://huggingface.co/datasets/ContextReq/Synthetic-Dataset-Childrens-Stories.melville-corpus
ContextLab Herman Melville Corpus
Dataset Description
This dataset contains works of Herman Melville (1819-1891), preprocessed for computational stylometry research. The texts were sourced from Project Gutenberg and cleaned for use in the paper "A Stylometric Application of Large Language Models" (Stropkay et al., 2025).
The corpus includes 10 books by Herman Melville, including Moby-Dick, Bartleby the Scrivener, and Typee. All text has been converted to lowercase and… See the full description on the dataset page: https://huggingface.co/datasets/contextlab/melville-corpus.baum-corpus
ContextLab L. Frank Baum Corpus
Dataset Description
This dataset contains works of L. Frank Baum (1856-1919), preprocessed for computational stylometry research. The texts were sourced from Project Gutenberg and cleaned for use in the paper "A Stylometric Application of Large Language Models" (Stropkay et al., 2025).
The corpus includes 14 books by L. Frank Baum, including The Wonderful Wizard of Oz series (14 books). All text has been converted to lowercase and cleaned… See the full description on the dataset page: https://huggingface.co/datasets/contextlab/baum-corpus.dickens-corpus
ContextLab Charles Dickens Corpus
Dataset Description
This dataset contains works of Charles Dickens (1812-1870), preprocessed for computational stylometry research. The texts were sourced from Project Gutenberg and cleaned for use in the paper "A Stylometric Application of Large Language Models" (Stropkay et al., 2025).
The corpus includes 14 books by Charles Dickens, including A Tale of Two Cities, Great Expectations, Oliver Twist, and David Copperfield. All text has… See the full description on the dataset page: https://huggingface.co/datasets/contextlab/dickens-corpus.long-context-retrieval-training-pool
Long context retrieval training pool
Long prompts with short, checkable answers. Each row is one complete message: a task instruction,
a long body of text that hides what the question is about, and the question itself, together with
every string an answer has to contain for it to be right. The bodies run from four thousand to
thirty-two thousand tokens. Three sources, laid out twice. Train on either layer or on
both.
pool.jsonl
Every source rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/long-context-retrieval-training-pool.in-context-grid-reasoning
In-Context Grid Reasoning (ICGR)
A small, fully synthetic benchmark for demonstration-conditioned rule induction:
each task shows 2–4 (input grid → output grid) support pairs that share one
hidden transformation, and the model must apply the same transformation to a
held-out query input.
It targets the same behaviour probed by recent in-context / latent-reasoning work
on ARC-AGI (e.g. BDH-CQ: In-Context Learning with Recurrent Latent Reasoning,
arXiv:2608.09888), but is… See the full description on the dataset page: https://huggingface.co/datasets/WhySoCodius/in-context-grid-reasoning.Organized_PreTrain_1k_Context#Plans11/Organized_PreTrain_1k_Context
A 1096-token hard-filtered, deduped, pretrain-ready merge of 11 Organized PreTrain datasets.
Built from Plans11 organized collections. Everything >1096 tokens was 100% trashed, never truncated. Global SHA256 dedup across all sources. Uploaded add-only (shard index computed from live Hub listing).
Destination: Plans11/Organized_PreTrain_1k_Context
Context Ceiling: 1096 tokens (cl100k_base proxy)
Total Kept: 2,713,413
Total Dropped >1096: 532,487
Total… See the full description on the dataset page: https://huggingface.co/datasets/11-47/Organized_PreTrain_1k_Context.task108_contextualabusedetection_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task108_contextualabusedetection_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task108_contextualabusedetection_classification.task270_csrg_counterfactual_context_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task270_csrg_counterfactual_context_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task270_csrg_counterfactual_context_generation.task966_ruletaker_fact_checking_based_on_given_context
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task966_ruletaker_fact_checking_based_on_given_context
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task966_ruletaker_fact_checking_based_on_given_context.kana-kanji-context
kana-kanji-context
Japanese kana-to-kanji conversion dataset with context for disambiguation.
Overview
Metric
Value
Total entries
77,277,970
File size
~7.4GB
Format
JSONL
Data Format
{
"input": "神経 [---]かがく",
"output": ["科学"],
"count": 1
}
{
"input": "この [---]さいご",
"output": ["最後", "最期"],
"count": 2
}
Fields
Field
Description
input
Context + [---] + reading (hiragana)
output
Correct kanji candidates (max… See the full description on the dataset page: https://huggingface.co/datasets/katsukiono/kana-kanji-context.task455_swag_context_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task455_swag_context_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task455_swag_context_generation.sql-create-context-instruction
Overview
This dataset is built upon SQL Create Context, which in turn was constructed using data from WikiSQL and Spider.
There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-SQL LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-SQL datasets. The CREATE TABLE statement can often be… See the full description on the dataset page: https://huggingface.co/datasets/bugdaryan/sql-create-context-instruction.sql-create-context-pt
Overview
Este dataset é uma versão traduzida para o português do dataset b-mc2/sql-create-context,
que foi construído a partir dos datasets WikiSQL e Spider. Ele contém exemplos de perguntas
em português, instruções SQL CREATE TABLE e consultas SQL que respondem às perguntas
utilizando a instrução CREATE TABLE como contexto.
O principal objetivo deste dataset é ajudar modelos de linguagem natural em português a gerar consultas
SQL precisas e contextualizadas, prevenindo a… See the full description on the dataset page: https://huggingface.co/datasets/emdemor/sql-create-context-pt.sql-create-context-copy
Fork of b-mc2/sql-create-context
Overview
This dataset builds from WikiSQL and Spider.
There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/sql-create-context-copy.wells-corpus
ContextLab H.G. Wells Corpus
Dataset Description
This dataset contains works of H.G. Wells (1866-1946), preprocessed for computational stylometry research. The texts were sourced from Project Gutenberg and cleaned for use in the paper "A Stylometric Application of Large Language Models" (Stropkay et al., 2025).
The corpus includes 12 books by H.G. Wells, including The Time Machine, The War of the Worlds, and The Invisible Man. All text has been converted to lowercase and… See the full description on the dataset page: https://huggingface.co/datasets/contextlab/wells-corpus.hausa-stem-reasoning-with-cultural-context
Hausa STEM Reasoning with Cultural Context
Abstract
We present the first large-scale bilingual Hausa-English STEM reasoning dataset with deep cultural adaptation, containing 2,640 high-quality question-answer pairs translated from the STEM-Reasoning-Complex dataset. Our work introduces the "Shehin Malamin Kimiyya" (The Wise Scholar of Science) translation framework, which transforms Western scientific concepts into culturally-embedded Hausa explanations using systematic… See the full description on the dataset page: https://huggingface.co/datasets/Tushe/hausa-stem-reasoning-with-cultural-context.ContextRL_Agentic
ContextRL-Agentic
The agentic (long-horizon) training set for ContextRL, used to train
ContextRL-Klear-AgentForge-8B
and ContextRL-Qwen3-8B-Agentic,
from the paper Context-Aware RL for Agentic and Multimodal LLMs.
Setup
Training and evaluation code, data construction pipelines, and detailed configurations are
available in the repository:
👉 https://github.com/xupy2003/ContextAwareRL
context-primitive-code-agent-pack-v0
Context Primitive Code-Agent Pack v0 — Free Funnel
Free product-specific instruction / Q&A seed material from Primitive Origins’ Context Primitive / Foundry tests.
This is a marketing / companion corpus for the Context Primitive stack — not a general public code-agent marketplace hero SKU.
What’s inside
JSONL splits under data/:
behavior_qa.train.jsonl / .eval.jsonl
instruction_test_generation.train.jsonl / .eval.jsonl
foundry/python_test_generation.*… See the full description on the dataset page: https://huggingface.co/datasets/Primitive-Origins/context-primitive-code-agent-pack-v0.kana-kanji-context
kana-kanji-context
Japanese kana-to-kanji conversion dataset with context for disambiguation.
Overview
Metric
Value
Total entries
77,277,970
File size
~7.4GB
Format
JSONL
Data Format
{
"input": "神経 [---]かがく",
"output": ["科学"],
"count": 1
}
{
"input": "この [---]さいご",
"output": ["最後", "最期"],
"count": 2
}
Fields
Field
Description
input
Context + [---] + reading (hiragana)
output
Correct kanji… See the full description on the dataset page: https://huggingface.co/datasets/0x3/kana-kanji-context.
