datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
long-context-qa-curated-20
Dataset Card / 数据集卡
Dataset Description / 数据集简介
This public release contains 20 curated samples selected from a 10,000-record long-context QA collection. It targets retrieval over long documents, cross-section evidence synthesis, numerical reasoning, timeline reconstruction, and structured answer evaluation. The public subset contains 15 short-answer questions and 5 multiple-choice questions, balanced across Chinese and English.
本公开版本从 10,000 条长上下文问答数据中精选 20… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/long-context-qa-curated-20.long_context_hin_22klong_context_hindi
Dataset
This dataset was filtered from AI4BHarat dataset sangraha,which is the largest high-quality, cleaned Indic language pretraining data containing 251B tokens summed up over 22 languages, extracted from curated sources, existing multilingual corpora and large scale translations.
This dataset contains only Hindi as of now
Information
First this dataset is mainly for long context training
The minimum len is 6000 and maximum len is 3754718
Getting started… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/long_context_hindi.LongMagpie_multidoc_longcontext_datasetlongcontext-haldetect
Long-Context Hallucination Detection Benchmark
A synthetic benchmark dataset for evaluating hallucination detection models on long documents (8K-24K tokens). This dataset is specifically designed to test models that can handle contexts beyond the typical 8K token limit.
Dataset Summary
Property
Value
Total samples
3,366
Token range
8,005 - 23,998
Average tokens
17,852
Hallucinated
1,681 (49.9%)
Supported
1,685 (50.1%)
Splits… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/longcontext-haldetect.Long-Context-Reasoning-Dataset
Description
본 데이터셋은 현재 대규모 언어 모델(LLM)이 장문 문서를 처리하고 복잡한 추론을 수행할 때 나타나는 핵심적인 한계를 보완하기 위해 구축되었습니다. 중국어, 영어, 한국어의 3개 언어로 구성된 총 7,500개의 고품질 학습 데이터를 포함하고 있습니다. 각 데이터는 장문의 텍스트를 기반으로 하며, 여러 문단과 문서에 걸쳐 정보를 종합하고 여러 단계의 논리적 추론 과정을 거쳐야 답변할 수 있는 질문으로 구성되어 있습니다. 본 데이터셋은 모델의 장거리 문맥 이해, 관련 정보 검색 및 추출, 논리적 추론 경로 구성, 근거 정보의 출처 추적 능력을 종합적이고 체계적으로 평가하는 데 활용할 수 있습니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/llm/2121?source=hf.kr
Specifications
Content
장문… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/Long-Context-Reasoning-Dataset.long-context-retrieval-training-pool
Long context retrieval training pool
Long prompts with short, checkable answers. Each row is one complete message: a task instruction,
a long body of text that hides what the question is about, and the question itself, together with
every string an answer has to contain for it to be right. The bodies run from four thousand to
thirty-two thousand tokens. Three sources, laid out twice. Train on either layer or on
both.
pool.jsonl
Every source rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/long-context-retrieval-training-pool.longcontext_datasetLongMagpie_singledoc_longcontext_dataset
LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions
This repository contains the code, models and datasets for our paper [LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions].
Quick Links
Overview
LongMagpie Models
LongMagpie Datasets
Datasets list
Train Llama-3-8B-LongMagpie-512K-Instruct
Requirements
Evaluation
Build your long-context instruction data
Bugs or Questions?
Overview… See the full description on the dataset page: https://huggingface.co/datasets/caskcsg/LongMagpie_singledoc_longcontext_dataset.mbpp-longcontext
MBPP Long-Context Dataset
Overview
MBPP Long-Context is a benchmark dataset that combines coding problems from the MBPP (Mostly Basic Python Problems) dataset with long-context distractors from BABILong. This dataset evaluates code generation performance under long-context conditions, testing whether models can maintain coding ability with stuffed context.
Dataset Structure
Data Fields
Each sample contains:
Original MBPP Fields… See the full description on the dataset page: https://huggingface.co/datasets/jannalu/mbpp-longcontext.long-context-probe-setslongcontext-cmrc2018-zhLongcontext-aozora-instruction長文用のinstructionデータセットです。
長文は以下の青空文庫データセットを利用しました。
globis-university/aozorabunko-clean
Limitation
このデータセットは、長文の質問応答スタイルを提示することを主な目的としています。質問応答の正誤についてのフィルタリングはあえて行っていません。
長文では一般に性能低下が認められるため困難なタスクとなります。フィルタリングすると困難なタスクのinstructionが消えてしまうためです。ファインチューニングで使用する場合は、チューニングする基盤モデルの性能によって、チューニング効果が大きく変わります。正答できるかどうかはモデルパラメータ、事前学習次第と考えられます。
License
CC BY 4.0
long_context_eval_set
long-context-text-summarization-alpaca-formatlong-context-qa-df
Dataset Card for "long-context-qa-df"
More Information needed
long-context-reasoning-v1LongContextCodeQA
LongContextCodeQA Java Dataset
Dataset Details
The file Java/questions_java.json contains the LongContextCodeQA Java dataset.
Each entry contains questions and multiple-choice options with the correct answer. The path for the code contexts for each context-length bucket (eg, 32k, 64k etc.) are also provided for each entry in the dataset.
Java repositories considered:
We considered 3 most starred public Java repositories to generate 85 questions -
Cassandra -… See the full description on the dataset page: https://huggingface.co/datasets/mjkishan/LongContextCodeQA.Long_Context_MT_ALL_EGlong-contextlong-context-baseline-bakeoff
Long-Context Data-Selection Bake-off — Shared Candidate Pool
The shared 16K candidate pool for comparing long-context data-selection methods on equal
footing. Every method (AttentionSpan, LongAttn, LongProc, ProLong, perplexity, ...) scores the
same 14,300 documents, picks its own top-800 under the same split, then trains
Llama-2-7B + 16K LoRA and evaluates on HELMET.
Files
File
Description
candidate_pool_16k_scored.parquet
The shared pool — 14,300 docs… See the full description on the dataset page: https://huggingface.co/datasets/KevinDavidHayes/long-context-baseline-bakeoff.long-context-llm-papers
Long-Context LLM Papers — FineSet
A research-paper dataset on Long-Context LLM Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on Long-Context LLM Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset
Quality-scored: quality_score… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/long-context-llm-papers.Longcontext-aozora-summary長文からの要約データセットです。
長文は以下の青空文庫データセットを利用しました。
globis-university/aozorabunko-clean
License
CC BY 4.0
longcontext-cmrc2018-zh-qrelsrankzephyr_longcontext_merged_140krankzephyr_longcontext_range80-100_merged_80klong-contextrankzephyr_longcontext_range80-100_merged_140kllama-2-7b-LongContext-mixed-32k-30APRIL2024longcontext-summarization-v1
