CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01b-mc2 /sql-create-context Overview This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/sql-create-context.texttext-generation10K<n<100K506 likes6.3k downloads3y agoHugging Face02yuyijiong /context_qa_sum_qwen3_synthetic Context-based QA and Summarization Synthetic Dataset Overview This dataset contains synthetic context-based question-answering (QA) and summarization data. The data was synthesized using: Source context: openbmb/Ultra-FineWeb Synthesis model: Qwen3-30B-A3B-Instruct-2507 Each context is obtained by taking the initial segment of raw pretraining text from Ultra-FineWeb, truncated to at most the corresponding number of tokens, while ensuring the truncation does not occur in… See the full description on the dataset page: https://huggingface.co/datasets/yuyijiong/context_qa_sum_qwen3_synthetic.texttext-generation10M<n<100M5 likes3.6k downloads6mo agoHugging Face03artefactory /ledger-long-context-KPI-QA LEDGER — Long-Context KPI Question Answering & Page Retrieval This dataset is part of the LEDGER (Long-context Evaluation of Documents for Grounded Extraction and Retrieval) benchmark. It supports two of the three LEDGER tasks: Page-level KPI retrieval — given a natural-language question about a financial KPI and the corresponding annual report, retrieve the relevant page(s). Each row includes TREC-style graded relevance judgments (qrels) over all candidate pages.… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-KPI-QA.tabularquestion-answering100K<n<1M14 likes3k downloads1mo agoHugging Face04contextecho2026 /persona-drift-contextecho ContextEcho — Released Dataset Per-cell evaluation corpus and donated session prefixes for the ContextEcho benchmark. This Hugging Face repository hosts the released dataset artifacts. The canonical project page, latest README, code, reproduction instructions, and donation workflow are maintained on GitHub: https://github.com/Accenture/ContextEcho Donate a coding-agent session: https://accenture.github.io/ContextEcho/donate/ For the formal datasheet, see DATASHEET.md.… See the full description on the dataset page: https://huggingface.co/datasets/contextecho2026/persona-drift-contextecho.text-generation10K<n<100K7 likes975 downloads3d agoHugging Face05TreeAILab /Multi-turn_Long-context_Benchmark_for_LLMs LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues Arxiv: https://www.arxiv.org/abs/2507.13681 Huggingface: https://huggingface.co/papers/2507.13681 Introduction LoopServe Multi-Turn Dialogue Benchmark is a comprehensive evaluation dataset comprising multiple diverse datasets designed to assess large language model performance in realistic conversational scenarios. Unlike traditional benchmarks that place queries only at the end… See the full description on the dataset page: https://huggingface.co/datasets/TreeAILab/Multi-turn_Long-context_Benchmark_for_LLMs.textquestion-answering1K<n<10K0 likes623 downloads1y agoHugging Face06ghostcc3 /mix-context-post-training-128k Mix-Context Post-Training Dataset for 128K Context Extension Overview Mix-Context Post-Training 128K is a dataset designed specifically for post-training context window extension of pretrained LLMs. It targets the stage after base pretraining, where a model is adapted to operate over much longer contexts (up to 128K tokens) while preserving short-context behavior. The dataset mixes short- and long-context packed sequences with a controlled length distribution to support:… See the full description on the dataset page: https://huggingface.co/datasets/ghostcc3/mix-context-post-training-128k.text-generation10K<n<100K3 likes563 downloads8mo agoHugging Face07huseyinatahaninan /ContextualIntegritySyntheticDataset Contextual Integrity Synthetic Dataset This repository contains the synthetic dataset introduced in the paper "Contextual Integrity in LLMs via Reasoning and Reinforcement Learning". Paper | Code | Blog Dataset Summary The Contextual Integrity (CI) synthetic dataset consists of 729 examples featuring diverse contexts and information disclosure norms. It is designed to instill reasoning capabilities in LLMs regarding what information is appropriate to share while… See the full description on the dataset page: https://huggingface.co/datasets/huseyinatahaninan/ContextualIntegritySyntheticDataset.texttext-generationn<1K2 likes388 downloads8mo agoHugging Face08contextlab /austen-corpus ContextLab Jane Austen Corpus Dataset Description This dataset contains works of Jane Austen (1775-1817), preprocessed for computational stylometry research. The texts were sourced from Project Gutenberg and cleaned for use in the paper "A Stylometric Application of Large Language Models" (Stropkay et al., 2025). The corpus includes 7 books by Jane Austen, including Pride and Prejudice, Sense and Sensibility, and Emma. All text has been converted to lowercase and cleaned… See the full description on the dataset page: https://huggingface.co/datasets/contextlab/austen-corpus.texttext-generationn<1K0 likes321 downloads11mo agoHugging Face09allenai /ContextEval Contextualized Evaluations: Taking the Guesswork Out of Language Model Evaluations Dataset Summary We provide here the data accompanying the paper: Contextualized Evaluations: Taking the Guesswork Out of Language Model Evaluations. Dataset Structure Data Instances We release the set of queries, as well as the autorater & human evaluation judgements collected for our experiments. Data overview List of queries: Data Structure The list… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ContextEval.text-generation10K<n<100K8 likes250 downloads2y agoHugging Face10damerajee /long_context_hindi Dataset This dataset was filtered from AI4BHarat dataset sangraha,which is the largest high-quality, cleaned Indic language pretraining data containing 251B tokens summed up over 22 languages, extracted from curated sources, existing multilingual corpora and large scale translations. This dataset contains only Hindi as of now Information First this dataset is mainly for long context training The minimum len is 6000 and maximum len is 3754718 Getting started… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/long_context_hindi.texttext-generation100K<n<1M1 likes220 downloads2y agoHugging Face11ContextReq /Synthetic-Dataset-Childrens-Stories**Status: released 13-09-2026, repacked 14-09-2026.** The 14-09-2026 repack replaced 58 items after the acceptance gates were strengthened (prompt-instruction leaks, markdown bullet lists and blockquotes); the other 29,942 are unchanged. Development stopped, pipeline released 17/09/26. SAMPLE RELEASE: 30,000 synthetic children's short stories for early-reader language modelling. Metrics Value genres 26 stories per genre 1.153-1.154K stories total characters 38… See the full description on the dataset page: https://huggingface.co/datasets/ContextReq/Synthetic-Dataset-Childrens-Stories.texttext-generation10K<n<100K1 likes210 downloads6d agoHugging Face12contextlab /melville-corpus ContextLab Herman Melville Corpus Dataset Description This dataset contains works of Herman Melville (1819-1891), preprocessed for computational stylometry research. The texts were sourced from Project Gutenberg and cleaned for use in the paper "A Stylometric Application of Large Language Models" (Stropkay et al., 2025). The corpus includes 10 books by Herman Melville, including Moby-Dick, Bartleby the Scrivener, and Typee. All text has been converted to lowercase and… See the full description on the dataset page: https://huggingface.co/datasets/contextlab/melville-corpus.texttext-generationn<1K0 likes154 downloads11mo agoHugging Face13contextlab /baum-corpus ContextLab L. Frank Baum Corpus Dataset Description This dataset contains works of L. Frank Baum (1856-1919), preprocessed for computational stylometry research. The texts were sourced from Project Gutenberg and cleaned for use in the paper "A Stylometric Application of Large Language Models" (Stropkay et al., 2025). The corpus includes 14 books by L. Frank Baum, including The Wonderful Wizard of Oz series (14 books). All text has been converted to lowercase and cleaned… See the full description on the dataset page: https://huggingface.co/datasets/contextlab/baum-corpus.texttext-generationn<1K1 likes149 downloads11mo agoHugging Face14contextlab /dickens-corpus ContextLab Charles Dickens Corpus Dataset Description This dataset contains works of Charles Dickens (1812-1870), preprocessed for computational stylometry research. The texts were sourced from Project Gutenberg and cleaned for use in the paper "A Stylometric Application of Large Language Models" (Stropkay et al., 2025). The corpus includes 14 books by Charles Dickens, including A Tale of Two Cities, Great Expectations, Oliver Twist, and David Copperfield. All text has… See the full description on the dataset page: https://huggingface.co/datasets/contextlab/dickens-corpus.texttext-generationn<1K0 likes127 downloads11mo agoHugging Face15Emulated-Inc /long-context-retrieval-training-pool Long context retrieval training pool Long prompts with short, checkable answers. Each row is one complete message: a task instruction, a long body of text that hides what the question is about, and the question itself, together with every string an answer has to contain for it to be right. The bodies run from four thousand to thirty-two thousand tokens. Three sources, laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/long-context-retrieval-training-pool.texttext-generation10K<n<100K1 likes125 downloads10d agoHugging Face16WhySoCodius /in-context-grid-reasoning In-Context Grid Reasoning (ICGR) A small, fully synthetic benchmark for demonstration-conditioned rule induction: each task shows 2–4 (input grid → output grid) support pairs that share one hidden transformation, and the model must apply the same transformation to a held-out query input. It targets the same behaviour probed by recent in-context / latent-reasoning work on ARC-AGI (e.g. BDH-CQ: In-Context Learning with Recurrent Latent Reasoning, arXiv:2608.09888), but is… See the full description on the dataset page: https://huggingface.co/datasets/WhySoCodius/in-context-grid-reasoning.tabulartext-generation1K<n<10K1 likes116 downloads19d agoHugging Face1711-47 /Organized_PreTrain_1k_Context#Plans11/Organized_PreTrain_1k_Context A 1096-token hard-filtered, deduped, pretrain-ready merge of 11 Organized PreTrain datasets. Built from Plans11 organized collections. Everything >1096 tokens was 100% trashed, never truncated. Global SHA256 dedup across all sources. Uploaded add-only (shard index computed from live Hub listing). Destination: Plans11/Organized_PreTrain_1k_Context Context Ceiling: 1096 tokens (cl100k_base proxy) Total Kept: 2,713,413 Total Dropped >1096: 532,487 Total… See the full description on the dataset page: https://huggingface.co/datasets/11-47/Organized_PreTrain_1k_Context.texttext-generation1M<n<10M0 likes112 downloads1mo agoHugging Face18Lots-of-LoRAs /task108_contextualabusedetection_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task108_contextualabusedetection_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task108_contextualabusedetection_classification.texttext-generation1K<n<10K0 likes102 downloads2y agoHugging Face19Lots-of-LoRAs /task270_csrg_counterfactual_context_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task270_csrg_counterfactual_context_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task270_csrg_counterfactual_context_generation.texttext-generation1K<n<10K0 likes95 downloads2y agoHugging Face20Lots-of-LoRAs /task966_ruletaker_fact_checking_based_on_given_context Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task966_ruletaker_fact_checking_based_on_given_context Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task966_ruletaker_fact_checking_based_on_given_context.texttext-generationn<1K0 likes92 downloads2y agoHugging Face21katsukiono /kana-kanji-context kana-kanji-context Japanese kana-to-kanji conversion dataset with context for disambiguation. Overview Metric Value Total entries 77,277,970 File size ~7.4GB Format JSONL Data Format { "input": "神経 [---]かがく", "output": ["科学"], "count": 1 } { "input": "この [---]さいご", "output": ["最後", "最期"], "count": 2 } Fields Field Description input Context + [---] + reading (hiragana) output Correct kanji candidates (max… See the full description on the dataset page: https://huggingface.co/datasets/katsukiono/kana-kanji-context.texttext-generation100M<n<1B1 likes91 downloads9mo agoHugging Face22Lots-of-LoRAs /task455_swag_context_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task455_swag_context_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task455_swag_context_generation.texttext-generation1K<n<10K0 likes75 downloads2y agoHugging Face23bugdaryan /sql-create-context-instruction Overview This dataset is built upon SQL Create Context, which in turn was constructed using data from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-SQL LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-SQL datasets. The CREATE TABLE statement can often be… See the full description on the dataset page: https://huggingface.co/datasets/bugdaryan/sql-create-context-instruction.texttext-generation10K<n<100K19 likes71 downloads3y agoHugging Face24emdemor /sql-create-context-pt Overview Este dataset é uma versão traduzida para o português do dataset b-mc2/sql-create-context, que foi construído a partir dos datasets WikiSQL e Spider. Ele contém exemplos de perguntas em português, instruções SQL CREATE TABLE e consultas SQL que respondem às perguntas utilizando a instrução CREATE TABLE como contexto. O principal objetivo deste dataset é ajudar modelos de linguagem natural em português a gerar consultas SQL precisas e contextualizadas, prevenindo a… See the full description on the dataset page: https://huggingface.co/datasets/emdemor/sql-create-context-pt.texttext-generation10K<n<100K2 likes69 downloads2y agoHugging Face25philschmid /sql-create-context-copy Fork of b-mc2/sql-create-context Overview This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/sql-create-context-copy.texttext-generation10K<n<100K4 likes68 downloads3y agoHugging Face26contextlab /wells-corpus ContextLab H.G. Wells Corpus Dataset Description This dataset contains works of H.G. Wells (1866-1946), preprocessed for computational stylometry research. The texts were sourced from Project Gutenberg and cleaned for use in the paper "A Stylometric Application of Large Language Models" (Stropkay et al., 2025). The corpus includes 12 books by H.G. Wells, including The Time Machine, The War of the Worlds, and The Invisible Man. All text has been converted to lowercase and… See the full description on the dataset page: https://huggingface.co/datasets/contextlab/wells-corpus.texttext-generationn<1K0 likes68 downloads11mo agoHugging Face27Tushe /hausa-stem-reasoning-with-cultural-context Hausa STEM Reasoning with Cultural Context Abstract We present the first large-scale bilingual Hausa-English STEM reasoning dataset with deep cultural adaptation, containing 2,640 high-quality question-answer pairs translated from the STEM-Reasoning-Complex dataset. Our work introduces the "Shehin Malamin Kimiyya" (The Wise Scholar of Science) translation framework, which transforms Western scientific concepts into culturally-embedded Hausa explanations using systematic… See the full description on the dataset page: https://huggingface.co/datasets/Tushe/hausa-stem-reasoning-with-cultural-context.textquestion-answering1K<n<10K1 likes68 downloads7mo agoHugging Face28xupy21 /ContextRL_Agentic ContextRL-Agentic The agentic (long-horizon) training set for ContextRL, used to train ContextRL-Klear-AgentForge-8B and ContextRL-Qwen3-8B-Agentic, from the paper Context-Aware RL for Agentic and Multimodal LLMs. Setup Training and evaluation code, data construction pipelines, and detailed configurations are available in the repository: 👉 https://github.com/xupy2003/ContextAwareRL texttext-generation1K<n<10K0 likes67 downloads3mo agoHugging Face29Primitive-Origins /context-primitive-code-agent-pack-v0 Context Primitive Code-Agent Pack v0 — Free Funnel Free product-specific instruction / Q&A seed material from Primitive Origins’ Context Primitive / Foundry tests. This is a marketing / companion corpus for the Context Primitive stack — not a general public code-agent marketplace hero SKU. What’s inside JSONL splits under data/: behavior_qa.train.jsonl / .eval.jsonl instruction_test_generation.train.jsonl / .eval.jsonl foundry/python_test_generation.*… See the full description on the dataset page: https://huggingface.co/datasets/Primitive-Origins/context-primitive-code-agent-pack-v0.texttext-generation1K<n<10K0 likes67 downloads4d agoHugging Face300x3 /kana-kanji-context kana-kanji-context Japanese kana-to-kanji conversion dataset with context for disambiguation. Overview Metric Value Total entries 77,277,970 File size ~7.4GB Format JSONL Data Format { "input": "神経 [---]かがく", "output": ["科学"], "count": 1 } { "input": "この [---]さいご", "output": ["最後", "最期"], "count": 2 } Fields Field Description input Context + [---] + reading (hiragana) output Correct kanji… See the full description on the dataset page: https://huggingface.co/datasets/0x3/kana-kanji-context.texttext-generation100M<n<1B0 likes56 downloads17d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.