CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01artefactory /ledger-long-context-KPI-QA LEDGER — Long-Context KPI Question Answering & Page Retrieval This dataset is part of the LEDGER (Long-context Evaluation of Documents for Grounded Extraction and Retrieval) benchmark. It supports two of the three LEDGER tasks: Page-level KPI retrieval — given a natural-language question about a financial KPI and the corresponding annual report, retrieve the relevant page(s). Each row includes TREC-style graded relevance judgments (qrels) over all candidate pages.… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-KPI-QA.tabularquestion-answering100K<n<1M14 likes2.9k downloads1mo agoHugging Face02artefactory /ledger-long-context-multi-kpi the LEDGER Long-Context Multi-KPI extraction datasets and benchmarks. OCR'd annual reports with ground-truth KPI values for financial information extraction benchmarking. Dataset Description This dataset pairs OCR-extracted annual report text (from DeepSeek OCR) with structured KPI ground-truth values. It is designed for evaluating LLM-based financial information extraction, retrieval, and needle-in-a-haystack tasks. Configs Config Reports… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-multi-kpi.imagetable-question-answering1K<n<10K13 likes2.7k downloads2mo agoHugging Face03yzhuang /Agentic-Long-Context-Understanding-QA 📖 Agentic Long Context Understanding 📖 Self-Taught Agentic Long Context Understanding (Arxiv). AgenticLU refines complex, long-context queries through self-clarifications and contextual grounding, enabling robust long-document understanding in a single pass. Installation Requirements This codebase is largely based on OpenRLHF and Helmet, kudos to them. The requirements are the same pip install openrlhf pip install -r ./HELMET/requirements.txt… See the full description on the dataset page: https://huggingface.co/datasets/yzhuang/Agentic-Long-Context-Understanding-QA.tabularquestion-answering100K<n<1M18 likes577 downloads1y agoHugging Face04placeholderlabs /pretrain-web-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 8,689,580,607 (8.7B) Trainable tokens 8,689,580,607 (8.7B) Documents 281,846 Shards 89 UTF-8 bytes 37,540,769,483 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-web-mix-long-context.tabular100K<n<1M0 likes469 downloads13d agoHugging Face05ymoslem /wmt-da-human-evaluation-long-context Dataset Summary Long-context / document-level dataset for Quality Estimation of Machine Translation. It is an augmented variant of the sentence-level WMT DA Human Evaluation dataset. In addition to individual sentences, it contains augmentations of 2, 4, 8, 16, and 32 sentences, among each language pair lp and domain. The raw column represents a weighted average of scores of augmented sentences using character lengths of src and mt as weights. The code used to apply the augmentation… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/wmt-da-human-evaluation-long-context.tabular1M<n<10M8 likes406 downloads2y agoHugging Face06placeholderlabs /pretrain-ultra-fineweb-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 1,367,358,024 (1.4B) Trainable tokens 1,367,358,024 (1.4B) Documents 48,077 Shards 73 UTF-8 bytes 6,386,740,105 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-ultra-fineweb-mix-long-context.tabular10K<n<100K1 likes380 downloads13d agoHugging Face07damerajee /long_context_hin_22ktext100K<n<1M0 likes375 downloads2y agoHugging Face08placeholderlabs /pretrain-commits-v2-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 12,571,681,749 (12.6B) Trainable tokens 4,460,160,435 (4.5B) Documents 992,475 Shards 327 UTF-8 bytes 49,288,867,997 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-commits-v2-mix-long-context.tabular1M<n<10M0 likes348 downloads8d agoHugging Face09nbtpj /multi-context-long-answer-datasettext1M<n<10M11 likes326 downloads4y agoHugging Face10damerajee /long_context_hindi Dataset This dataset was filtered from AI4BHarat dataset sangraha,which is the largest high-quality, cleaned Indic language pretraining data containing 251B tokens summed up over 22 languages, extracted from curated sources, existing multilingual corpora and large scale translations. This dataset contains only Hindi as of now Information First this dataset is mainly for long context training The minimum len is 6000 and maximum len is 3754718 Getting started… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/long_context_hindi.texttext-generation100K<n<1M1 likes222 downloads2y agoHugging Face11placeholderlabs /pretrain-academic-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 43,694,042,993 (43.7B) Trainable tokens 43,694,042,993 (43.7B) Documents 1,001,557 Shards 373 UTF-8 bytes 183,279,720,921 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-academic-mix-long-context.tabular1M<n<10M0 likes161 downloads13d agoHugging Face12Seerkfang /LongMagpie_multidoc_longcontext_datasettext100K<n<1M5 likes159 downloads1y agoHugging Face13llm-semantic-router /longcontext-haldetect Long-Context Hallucination Detection Benchmark A synthetic benchmark dataset for evaluating hallucination detection models on long documents (8K-24K tokens). This dataset is specifically designed to test models that can handle contexts beyond the typical 8K token limit. Dataset Summary Property Value Total samples 3,366 Token range 8,005 - 23,998 Average tokens 17,852 Hallucinated 1,681 (49.9%) Supported 1,685 (50.1%) Splits… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/longcontext-haldetect.texttoken-classificationn<1K0 likes150 downloads9mo agoHugging Face14Emulated-Inc /long-context-retrieval-training-pool Long context retrieval training pool Long prompts with short, checkable answers. Each row is one complete message: a task instruction, a long body of text that hides what the question is about, and the question itself, together with every string an answer has to contain for it to be right. The bodies run from four thousand to thirty-two thousand tokens. Three sources, laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/long-context-retrieval-training-pool.texttext-generation10K<n<100K1 likes126 downloads12d agoHugging Face15crellis /longcontext_datasettext1M<n<10M0 likes118 downloads5mo agoHugging Face16placeholderlabs /pretrain-encyclopedic-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 587,625,128 (587.6M) Trainable tokens 587,625,128 (587.6M) Documents 23,631 Shards 9 UTF-8 bytes 1,978,753,989 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-encyclopedic-mix-long-context.tabular10K<n<100K1 likes102 downloads13d agoHugging Face17placeholderlabs /pretrain-repository-v2-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 130,158,375,824 (130.2B) Trainable tokens 130,158,375,824 (130.2B) Documents 2,578,578 Shards 1,168 UTF-8 bytes 535,260,344,241 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-repository-v2-mix-long-context.tabular1M<n<10M0 likes89 downloads8d agoHugging Face18VivekChauhan06 /repobench_python_long_context RepoBench v1.1 (Python) Introduction This dataset presents the Python portion of RepoBench v1.1 (ICLR 2024). The data encompasses a collection from GitHub, spanning the period from October 6th to December 31st, 2023. With a commitment to data integrity, a deduplication process based on file content has been implemented against the Stack v2 dataset (coming soon), aiming to mitigate data leakage and memorization concerns. Features The dataset provides the… See the full description on the dataset page: https://huggingface.co/datasets/VivekChauhan06/repobench_python_long_context.tabular10K<n<100K1 likes70 downloads2y agoHugging Face19jinaai /longcontext-cmrc2018-zhtext1K<n<10K2 likes67 downloads3y agoHugging Face20placeholderlabs /pretrain-nemotron-math-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 1,446,296,439 (1.4B) Trainable tokens 1,446,296,439 (1.4B) Documents 42,379 Shards 23 UTF-8 bytes 4,965,563,314 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-nemotron-math-mix-long-context.tabular10K<n<100K0 likes64 downloads13d agoHugging Face21antash420 /long-context-text-summarization-alpaca-formattext100K<n<1M1 likes44 downloads2y agoHugging Face22Abzu /long-context-qa-df Dataset Card for "long-context-qa-df" More Information needed textn<1K2 likes43 downloads3y agoHugging Face23AIGym /long-context-reasoning-v1text10K<n<100K0 likes39 downloads1y agoHugging Face24RanaGaber /Long_Context_MT_ALL_EGtext10K<n<100K0 likes34 downloads2mo agoHugging Face25sci-datasets /sci-long-contexttext100K<n<1M0 likes31 downloads7mo agoHugging Face26pritamdeb68 /FineWeb-long-context-documentstabular100K<n<1M1 likes29 downloads1y agoHugging Face27Kyle1668 /sfm-midtraining-mix-dclm-long-context-passages-blocklist-filteredtabular10K<n<100K0 likes29 downloads10mo agoHugging Face28jaehyeokdoo2 /Qwen2.5-32B-Instruct_short_long_context_140ktext100K<n<1M0 likes27 downloads2y agoHugging Face29jaehyeokdoo2 /Qwen2.5-32B-Instruct_long_context_range80-100_train_data_40ktext10K<n<100K0 likes26 downloads2y agoHugging Face30jinaai /longcontext-cmrc2018-zh-qrelstext1K<n<10K1 likes24 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.