CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01b-mc2 /sql-create-context Overview This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/sql-create-context.texttext-generation10K<n<100K506 likes6.3k downloads3y agoHugging Face02yuyijiong /context_qa_sum_qwen3_synthetic Context-based QA and Summarization Synthetic Dataset Overview This dataset contains synthetic context-based question-answering (QA) and summarization data. The data was synthesized using: Source context: openbmb/Ultra-FineWeb Synthesis model: Qwen3-30B-A3B-Instruct-2507 Each context is obtained by taking the initial segment of raw pretraining text from Ultra-FineWeb, truncated to at most the corresponding number of tokens, while ensuring the truncation does not occur in… See the full description on the dataset page: https://huggingface.co/datasets/yuyijiong/context_qa_sum_qwen3_synthetic.texttext-generation10M<n<100M5 likes3.7k downloads6mo agoHugging Face03artefactory /ledger-long-context-KPI-QA LEDGER — Long-Context KPI Question Answering & Page Retrieval This dataset is part of the LEDGER (Long-context Evaluation of Documents for Grounded Extraction and Retrieval) benchmark. It supports two of the three LEDGER tasks: Page-level KPI retrieval — given a natural-language question about a financial KPI and the corresponding annual report, retrieve the relevant page(s). Each row includes TREC-style graded relevance judgments (qrels) over all candidate pages.… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-KPI-QA.tabularquestion-answering100K<n<1M14 likes3k downloads1mo agoHugging Face04TreeAILab /Multi-turn_Long-context_Benchmark_for_LLMs LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues Arxiv: https://www.arxiv.org/abs/2507.13681 Huggingface: https://huggingface.co/papers/2507.13681 Introduction LoopServe Multi-Turn Dialogue Benchmark is a comprehensive evaluation dataset comprising multiple diverse datasets designed to assess large language model performance in realistic conversational scenarios. Unlike traditional benchmarks that place queries only at the end… See the full description on the dataset page: https://huggingface.co/datasets/TreeAILab/Multi-turn_Long-context_Benchmark_for_LLMs.textquestion-answering1K<n<10K0 likes487 downloads1y agoHugging Face05huseyinatahaninan /ContextualIntegritySyntheticDataset Contextual Integrity Synthetic Dataset This repository contains the synthetic dataset introduced in the paper "Contextual Integrity in LLMs via Reasoning and Reinforcement Learning". Paper | Code | Blog Dataset Summary The Contextual Integrity (CI) synthetic dataset consists of 729 examples featuring diverse contexts and information disclosure norms. It is designed to instill reasoning capabilities in LLMs regarding what information is appropriate to share while… See the full description on the dataset page: https://huggingface.co/datasets/huseyinatahaninan/ContextualIntegritySyntheticDataset.texttext-generationn<1K2 likes419 downloads8mo agoHugging Face06contextlab /austen-corpus ContextLab Jane Austen Corpus Dataset Description This dataset contains works of Jane Austen (1775-1817), preprocessed for computational stylometry research. The texts were sourced from Project Gutenberg and cleaned for use in the paper "A Stylometric Application of Large Language Models" (Stropkay et al., 2025). The corpus includes 7 books by Jane Austen, including Pride and Prejudice, Sense and Sensibility, and Emma. All text has been converted to lowercase and cleaned… See the full description on the dataset page: https://huggingface.co/datasets/contextlab/austen-corpus.texttext-generationn<1K0 likes321 downloads11mo agoHugging Face07damerajee /long_context_hindi Dataset This dataset was filtered from AI4BHarat dataset sangraha,which is the largest high-quality, cleaned Indic language pretraining data containing 251B tokens summed up over 22 languages, extracted from curated sources, existing multilingual corpora and large scale translations. This dataset contains only Hindi as of now Information First this dataset is mainly for long context training The minimum len is 6000 and maximum len is 3754718 Getting started… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/long_context_hindi.texttext-generation100K<n<1M1 likes222 downloads2y agoHugging Face08ContextReq /Synthetic-Dataset-Childrens-Stories**Status: released 13-09-2026, repacked 14-09-2026.** The 14-09-2026 repack replaced 58 items after the acceptance gates were strengthened (prompt-instruction leaks, markdown bullet lists and blockquotes); the other 29,942 are unchanged. Development stopped, pipeline released 17/09/26. SAMPLE RELEASE: 30,000 synthetic children's short stories for early-reader language modelling. Metrics Value genres 26 stories per genre 1.153-1.154K stories total characters 38… See the full description on the dataset page: https://huggingface.co/datasets/ContextReq/Synthetic-Dataset-Childrens-Stories.texttext-generation10K<n<100K1 likes210 downloads7d agoHugging Face09contextlab /melville-corpus ContextLab Herman Melville Corpus Dataset Description This dataset contains works of Herman Melville (1819-1891), preprocessed for computational stylometry research. The texts were sourced from Project Gutenberg and cleaned for use in the paper "A Stylometric Application of Large Language Models" (Stropkay et al., 2025). The corpus includes 10 books by Herman Melville, including Moby-Dick, Bartleby the Scrivener, and Typee. All text has been converted to lowercase and… See the full description on the dataset page: https://huggingface.co/datasets/contextlab/melville-corpus.texttext-generationn<1K0 likes154 downloads11mo agoHugging Face10contextlab /baum-corpus ContextLab L. Frank Baum Corpus Dataset Description This dataset contains works of L. Frank Baum (1856-1919), preprocessed for computational stylometry research. The texts were sourced from Project Gutenberg and cleaned for use in the paper "A Stylometric Application of Large Language Models" (Stropkay et al., 2025). The corpus includes 14 books by L. Frank Baum, including The Wonderful Wizard of Oz series (14 books). All text has been converted to lowercase and cleaned… See the full description on the dataset page: https://huggingface.co/datasets/contextlab/baum-corpus.texttext-generationn<1K1 likes145 downloads11mo agoHugging Face11contextlab /dickens-corpus ContextLab Charles Dickens Corpus Dataset Description This dataset contains works of Charles Dickens (1812-1870), preprocessed for computational stylometry research. The texts were sourced from Project Gutenberg and cleaned for use in the paper "A Stylometric Application of Large Language Models" (Stropkay et al., 2025). The corpus includes 14 books by Charles Dickens, including A Tale of Two Cities, Great Expectations, Oliver Twist, and David Copperfield. All text has… See the full description on the dataset page: https://huggingface.co/datasets/contextlab/dickens-corpus.texttext-generationn<1K0 likes129 downloads11mo agoHugging Face12Emulated-Inc /long-context-retrieval-training-pool Long context retrieval training pool Long prompts with short, checkable answers. Each row is one complete message: a task instruction, a long body of text that hides what the question is about, and the question itself, together with every string an answer has to contain for it to be right. The bodies run from four thousand to thirty-two thousand tokens. Three sources, laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/long-context-retrieval-training-pool.texttext-generation10K<n<100K1 likes125 downloads11d agoHugging Face13WhySoCodius /in-context-grid-reasoning In-Context Grid Reasoning (ICGR) A small, fully synthetic benchmark for demonstration-conditioned rule induction: each task shows 2–4 (input grid → output grid) support pairs that share one hidden transformation, and the model must apply the same transformation to a held-out query input. It targets the same behaviour probed by recent in-context / latent-reasoning work on ARC-AGI (e.g. BDH-CQ: In-Context Learning with Recurrent Latent Reasoning, arXiv:2608.09888), but is… See the full description on the dataset page: https://huggingface.co/datasets/WhySoCodius/in-context-grid-reasoning.tabulartext-generation1K<n<10K1 likes122 downloads20d agoHugging Face1411-47 /Organized_PreTrain_1k_Context#Plans11/Organized_PreTrain_1k_Context A 1096-token hard-filtered, deduped, pretrain-ready merge of 11 Organized PreTrain datasets. Built from Plans11 organized collections. Everything >1096 tokens was 100% trashed, never truncated. Global SHA256 dedup across all sources. Uploaded add-only (shard index computed from live Hub listing). Destination: Plans11/Organized_PreTrain_1k_Context Context Ceiling: 1096 tokens (cl100k_base proxy) Total Kept: 2,713,413 Total Dropped >1096: 532,487 Total… See the full description on the dataset page: https://huggingface.co/datasets/11-47/Organized_PreTrain_1k_Context.texttext-generation1M<n<10M0 likes113 downloads1mo agoHugging Face15Lots-of-LoRAs /task108_contextualabusedetection_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task108_contextualabusedetection_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task108_contextualabusedetection_classification.texttext-generation1K<n<10K0 likes100 downloads2y agoHugging Face16katsukiono /kana-kanji-context kana-kanji-context Japanese kana-to-kanji conversion dataset with context for disambiguation. Overview Metric Value Total entries 77,277,970 File size ~7.4GB Format JSONL Data Format { "input": "神経 [---]かがく", "output": ["科学"], "count": 1 } { "input": "この [---]さいご", "output": ["最後", "最期"], "count": 2 } Fields Field Description input Context + [---] + reading (hiragana) output Correct kanji candidates (max… See the full description on the dataset page: https://huggingface.co/datasets/katsukiono/kana-kanji-context.texttext-generation100M<n<1B1 likes93 downloads9mo agoHugging Face17Lots-of-LoRAs /task270_csrg_counterfactual_context_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task270_csrg_counterfactual_context_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task270_csrg_counterfactual_context_generation.texttext-generation1K<n<10K0 likes90 downloads2y agoHugging Face18Lots-of-LoRAs /task966_ruletaker_fact_checking_based_on_given_context Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task966_ruletaker_fact_checking_based_on_given_context Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task966_ruletaker_fact_checking_based_on_given_context.texttext-generationn<1K0 likes88 downloads2y agoHugging Face19Lots-of-LoRAs /task455_swag_context_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task455_swag_context_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task455_swag_context_generation.texttext-generation1K<n<10K0 likes69 downloads2y agoHugging Face20contextlab /wells-corpus ContextLab H.G. Wells Corpus Dataset Description This dataset contains works of H.G. Wells (1866-1946), preprocessed for computational stylometry research. The texts were sourced from Project Gutenberg and cleaned for use in the paper "A Stylometric Application of Large Language Models" (Stropkay et al., 2025). The corpus includes 12 books by H.G. Wells, including The Time Machine, The War of the Worlds, and The Invisible Man. All text has been converted to lowercase and… See the full description on the dataset page: https://huggingface.co/datasets/contextlab/wells-corpus.texttext-generationn<1K0 likes68 downloads11mo agoHugging Face21bugdaryan /sql-create-context-instruction Overview This dataset is built upon SQL Create Context, which in turn was constructed using data from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-SQL LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-SQL datasets. The CREATE TABLE statement can often be… See the full description on the dataset page: https://huggingface.co/datasets/bugdaryan/sql-create-context-instruction.texttext-generation10K<n<100K19 likes67 downloads3y agoHugging Face22Tushe /hausa-stem-reasoning-with-cultural-context Hausa STEM Reasoning with Cultural Context Abstract We present the first large-scale bilingual Hausa-English STEM reasoning dataset with deep cultural adaptation, containing 2,640 high-quality question-answer pairs translated from the STEM-Reasoning-Complex dataset. Our work introduces the "Shehin Malamin Kimiyya" (The Wise Scholar of Science) translation framework, which transforms Western scientific concepts into culturally-embedded Hausa explanations using systematic… See the full description on the dataset page: https://huggingface.co/datasets/Tushe/hausa-stem-reasoning-with-cultural-context.textquestion-answering1K<n<10K1 likes67 downloads7mo agoHugging Face23Primitive-Origins /context-primitive-code-agent-pack-v0 Context Primitive Code-Agent Pack v0 — Free Funnel Free product-specific instruction / Q&A seed material from Primitive Origins’ Context Primitive / Foundry tests. This is a marketing / companion corpus for the Context Primitive stack — not a general public code-agent marketplace hero SKU. What’s inside JSONL splits under data/: behavior_qa.train.jsonl / .eval.jsonl instruction_test_generation.train.jsonl / .eval.jsonl foundry/python_test_generation.*… See the full description on the dataset page: https://huggingface.co/datasets/Primitive-Origins/context-primitive-code-agent-pack-v0.texttext-generation1K<n<10K0 likes67 downloads5d agoHugging Face24philschmid /sql-create-context-copy Fork of b-mc2/sql-create-context Overview This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/sql-create-context-copy.texttext-generation10K<n<100K4 likes66 downloads3y agoHugging Face25emdemor /sql-create-context-pt Overview Este dataset é uma versão traduzida para o português do dataset b-mc2/sql-create-context, que foi construído a partir dos datasets WikiSQL e Spider. Ele contém exemplos de perguntas em português, instruções SQL CREATE TABLE e consultas SQL que respondem às perguntas utilizando a instrução CREATE TABLE como contexto. O principal objetivo deste dataset é ajudar modelos de linguagem natural em português a gerar consultas SQL precisas e contextualizadas, prevenindo a… See the full description on the dataset page: https://huggingface.co/datasets/emdemor/sql-create-context-pt.texttext-generation10K<n<100K2 likes66 downloads2y agoHugging Face26detakarang /sql-create-context-id Overview This dataset is a fork from sql-create-context This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from… See the full description on the dataset page: https://huggingface.co/datasets/detakarang/sql-create-context-id.texttext-generation10K<n<100K0 likes56 downloads3y agoHugging Face270x3 /kana-kanji-context kana-kanji-context Japanese kana-to-kanji conversion dataset with context for disambiguation. Overview Metric Value Total entries 77,277,970 File size ~7.4GB Format JSONL Data Format { "input": "神経 [---]かがく", "output": ["科学"], "count": 1 } { "input": "この [---]さいご", "output": ["最後", "最期"], "count": 2 } Fields Field Description input Context + [---] + reading (hiragana) output Correct kanji… See the full description on the dataset page: https://huggingface.co/datasets/0x3/kana-kanji-context.texttext-generation100M<n<1B0 likes56 downloads18d agoHugging Face28hiltch /pandas-create-context Overview This dataset is built from sql-create-context, which in itself builds from WikiSQL and Spider. I have used GPT4 to translate the SQL schema into pandas DataFrame schem initialization statements and to translate the SQL queries into pandas queries. There are 862 examples of natural language queries, pandas DataFrame creation statements, and pandas query answering the question using the DataFrame creation statement as context. This dataset was built with text-to-pandas… See the full description on the dataset page: https://huggingface.co/datasets/hiltch/pandas-create-context.texttext-generation10K<n<100K2 likes55 downloads3y agoHugging Face29portkey /truthful_qa_context Dataset Card for truthful_qa_context Dataset Summary TruthfulQA Context is an extension of the TruthfulQA benchmark, specifically designed to enhance its utility for models that rely on Retrieval-Augmented Generation (RAG). This version includes the original questions and answers from TruthfulQA, along with the added context text directly associated with each question. This additional context aims to provide immediate reference material for models, making it particularly… See the full description on the dataset page: https://huggingface.co/datasets/portkey/truthful_qa_context.texttext-generationn<1K8 likes53 downloads3y agoHugging Face30xupy21 /ContextRL_Agentic ContextRL-Agentic The agentic (long-horizon) training set for ContextRL, used to train ContextRL-Klear-AgentForge-8B and ContextRL-Qwen3-8B-Agentic, from the paper Context-Aware RL for Agentic and Multimodal LLMs. Setup Training and evaluation code, data construction pipelines, and detailed configurations are available in the repository: 👉 https://github.com/xupy2003/ContextAwareRL texttext-generation1K<n<10K0 likes51 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.