CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LianeMarilin /long-context-qa-curated-20 Dataset Card / 数据集卡 Dataset Description / 数据集简介 This public release contains 20 curated samples selected from a 10,000-record long-context QA collection. It targets retrieval over long documents, cross-section evidence synthesis, numerical reasoning, timeline reconstruction, and structured answer evaluation. The public subset contains 15 short-answer questions and 5 multiple-choice questions, balanced across Chinese and English. 本公开版本从 10,000 条长上下文问答数据中精选 20… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/long-context-qa-curated-20.textquestion-answeringn<1K0 likes647 downloads23d agoHugging Face02damerajee /long_context_hin_22ktext100K<n<1M0 likes375 downloads2y agoHugging Face03damerajee /long_context_hindi Dataset This dataset was filtered from AI4BHarat dataset sangraha,which is the largest high-quality, cleaned Indic language pretraining data containing 251B tokens summed up over 22 languages, extracted from curated sources, existing multilingual corpora and large scale translations. This dataset contains only Hindi as of now Information First this dataset is mainly for long context training The minimum len is 6000 and maximum len is 3754718 Getting started… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/long_context_hindi.texttext-generation100K<n<1M1 likes222 downloads2y agoHugging Face04Seerkfang /LongMagpie_multidoc_longcontext_datasettext100K<n<1M5 likes159 downloads1y agoHugging Face05llm-semantic-router /longcontext-haldetect Long-Context Hallucination Detection Benchmark A synthetic benchmark dataset for evaluating hallucination detection models on long documents (8K-24K tokens). This dataset is specifically designed to test models that can handle contexts beyond the typical 8K token limit. Dataset Summary Property Value Total samples 3,366 Token range 8,005 - 23,998 Average tokens 17,852 Hallucinated 1,681 (49.9%) Supported 1,685 (50.1%) Splits… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/longcontext-haldetect.texttoken-classificationn<1K0 likes150 downloads9mo agoHugging Face06Nexdata-kr /Long-Context-Reasoning-Dataset Description 본 데이터셋은 현재 대규모 언어 모델(LLM)이 장문 문서를 처리하고 복잡한 추론을 수행할 때 나타나는 핵심적인 한계를 보완하기 위해 구축되었습니다. 중국어, 영어, 한국어의 3개 언어로 구성된 총 7,500개의 고품질 학습 데이터를 포함하고 있습니다. 각 데이터는 장문의 텍스트를 기반으로 하며, 여러 문단과 문서에 걸쳐 정보를 종합하고 여러 단계의 논리적 추론 과정을 거쳐야 답변할 수 있는 질문으로 구성되어 있습니다. 본 데이터셋은 모델의 장거리 문맥 이해, 관련 정보 검색 및 추출, 논리적 추론 경로 구성, 근거 정보의 출처 추적 능력을 종합적이고 체계적으로 평가하는 데 활용할 수 있습니다. 자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/llm/2121?source=hf.kr Specifications Content 장문… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/Long-Context-Reasoning-Dataset.imagen<1K0 likes140 downloads16d agoHugging Face07Emulated-Inc /long-context-retrieval-training-pool Long context retrieval training pool Long prompts with short, checkable answers. Each row is one complete message: a task instruction, a long body of text that hides what the question is about, and the question itself, together with every string an answer has to contain for it to be right. The bodies run from four thousand to thirty-two thousand tokens. Three sources, laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/long-context-retrieval-training-pool.texttext-generation10K<n<100K1 likes126 downloads13d agoHugging Face08crellis /longcontext_datasettext1M<n<10M0 likes118 downloads5mo agoHugging Face09caskcsg /LongMagpie_singledoc_longcontext_dataset LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions This repository contains the code, models and datasets for our paper [LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions]. Quick Links Overview LongMagpie Models LongMagpie Datasets Datasets list Train Llama-3-8B-LongMagpie-512K-Instruct Requirements Evaluation Build your long-context instruction data Bugs or Questions? Overview… See the full description on the dataset page: https://huggingface.co/datasets/caskcsg/LongMagpie_singledoc_longcontext_dataset.text100K<n<1M5 likes105 downloads1y agoHugging Face10jannalu /mbpp-longcontext MBPP Long-Context Dataset Overview MBPP Long-Context is a benchmark dataset that combines coding problems from the MBPP (Mostly Basic Python Problems) dataset with long-context distractors from BABILong. This dataset evaluates code generation performance under long-context conditions, testing whether models can maintain coding ability with stuffed context. Dataset Structure Data Fields Each sample contains: Original MBPP Fields… See the full description on the dataset page: https://huggingface.co/datasets/jannalu/mbpp-longcontext.tabulartext-generation10K<n<100K0 likes91 downloads11mo agoHugging Face11yuzhaouoe /long-context-probe-setstext10K<n<100K0 likes75 downloads11d agoHugging Face12jinaai /longcontext-cmrc2018-zhtext1K<n<10K2 likes67 downloads3y agoHugging Face13aixsatoshi /Longcontext-aozora-instruction長文用のinstructionデータセットです。 長文は以下の青空文庫データセットを利用しました。 globis-university/aozorabunko-clean Limitation このデータセットは、長文の質問応答スタイルを提示することを主な目的としています。質問応答の正誤についてのフィルタリングはあえて行っていません。 長文では一般に性能低下が認められるため困難なタスクとなります。フィルタリングすると困難なタスクのinstructionが消えてしまうためです。ファインチューニングで使用する場合は、チューニングする基盤モデルの性能によって、チューニング効果が大きく変わります。正答できるかどうかはモデルパラメータ、事前学習次第と考えられます。 License CC BY 4.0 tabular1K<n<10K9 likes50 downloads2y agoHugging Face14baseten /long_context_eval_set textn<1K0 likes45 downloads2y agoHugging Face15antash420 /long-context-text-summarization-alpaca-formattext100K<n<1M1 likes44 downloads2y agoHugging Face16Abzu /long-context-qa-df Dataset Card for "long-context-qa-df" More Information needed textn<1K2 likes43 downloads3y agoHugging Face17AIGym /long-context-reasoning-v1text10K<n<100K0 likes39 downloads1y agoHugging Face18mjkishan /LongContextCodeQA LongContextCodeQA Java Dataset Dataset Details The file Java/questions_java.json contains the LongContextCodeQA Java dataset. Each entry contains questions and multiple-choice options with the correct answer. The path for the code contexts for each context-length bucket (eg, 32k, 64k etc.) are also provided for each entry in the dataset. Java repositories considered: We considered 3 most starred public Java repositories to generate 85 questions - Cassandra -… See the full description on the dataset page: https://huggingface.co/datasets/mjkishan/LongContextCodeQA.text0 likes36 downloads7mo agoHugging Face19RanaGaber /Long_Context_MT_ALL_EGtext10K<n<100K0 likes34 downloads2mo agoHugging Face20lightseekorg /long-contexttextn<1K0 likes33 downloads4mo agoHugging Face21KevinDavidHayes /long-context-baseline-bakeoff Long-Context Data-Selection Bake-off — Shared Candidate Pool The shared 16K candidate pool for comparing long-context data-selection methods on equal footing. Every method (AttentionSpan, LongAttn, LongProc, ProLong, perplexity, ...) scores the same 14,300 documents, picks its own top-800 under the same split, then trains Llama-2-7B + 16K LoRA and evaluates on HELMET. Files File Description candidate_pool_16k_scored.parquet The shared pool — 14,300 docs… See the full description on the dataset page: https://huggingface.co/datasets/KevinDavidHayes/long-context-baseline-bakeoff.texttext-generation1K<n<10K0 likes33 downloads2mo agoHugging Face22fineset-io /long-context-llm-papers Long-Context LLM Papers — FineSet A research-paper dataset on Long-Context LLM Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on Long-Context LLM Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset Quality-scored: quality_score… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/long-context-llm-papers.tabulartext-classificationn<1K0 likes32 downloads3mo agoHugging Face23aixsatoshi /Longcontext-aozora-summary長文からの要約データセットです。 長文は以下の青空文庫データセットを利用しました。 globis-university/aozorabunko-clean License CC BY 4.0 text1K<n<10K7 likes28 downloads2y agoHugging Face24jinaai /longcontext-cmrc2018-zh-qrelstext1K<n<10K1 likes24 downloads3y agoHugging Face25jaehyeokdoo2 /rankzephyr_longcontext_merged_140ktext100K<n<1M0 likes23 downloads2y agoHugging Face26jaehyeokdoo2 /rankzephyr_longcontext_range80-100_merged_80ktext10K<n<100K0 likes23 downloads2y agoHugging Face27tilde-research /long-contexttabularn<1K1 likes23 downloads1y agoHugging Face28jaehyeokdoo2 /rankzephyr_longcontext_range80-100_merged_140ktext100K<n<1M0 likes21 downloads2y agoHugging Face29thusinh1969 /llama-2-7b-LongContext-mixed-32k-30APRIL2024text10K<n<100K0 likes20 downloads2y agoHugging Face30AIGym /longcontext-summarization-v1text100K<n<1M0 likes20 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.