CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01artefactory /ledger-long-context-multi-kpi the LEDGER Long-Context Multi-KPI extraction datasets and benchmarks. OCR'd annual reports with ground-truth KPI values for financial information extraction benchmarking. Dataset Description This dataset pairs OCR-extracted annual report text (from DeepSeek OCR) with structured KPI ground-truth values. It is designed for evaluating LLM-based financial information extraction, retrieval, and needle-in-a-haystack tasks. Configs Config Reports… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-multi-kpi.imagetable-question-answering1K<n<10K12 likes3k downloads2mo agoHugging Face02artefactory /ledger-long-context-KPI-QA LEDGER — Long-Context KPI Question Answering & Page Retrieval This dataset is part of the LEDGER (Long-context Evaluation of Documents for Grounded Extraction and Retrieval) benchmark. It supports two of the three LEDGER tasks: Page-level KPI retrieval — given a natural-language question about a financial KPI and the corresponding annual report, retrieve the relevant page(s). Each row includes TREC-style graded relevance judgments (qrels) over all candidate pages.… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-KPI-QA.tabularquestion-answering100K<n<1M14 likes3k downloads1mo agoHugging Face03yzhuang /Agentic-Long-Context-Understanding-QA 📖 Agentic Long Context Understanding 📖 Self-Taught Agentic Long Context Understanding (Arxiv). AgenticLU refines complex, long-context queries through self-clarifications and contextual grounding, enabling robust long-document understanding in a single pass. Installation Requirements This codebase is largely based on OpenRLHF and Helmet, kudos to them. The requirements are the same pip install openrlhf pip install -r ./HELMET/requirements.txt… See the full description on the dataset page: https://huggingface.co/datasets/yzhuang/Agentic-Long-Context-Understanding-QA.tabularquestion-answering100K<n<1M17 likes562 downloads1y agoHugging Face04placeholderlabs /pretrain-web-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 8,689,580,607 (8.7B) Trainable tokens 8,689,580,607 (8.7B) Documents 281,846 Shards 89 UTF-8 bytes 37,540,769,483 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-web-mix-long-context.tabular100K<n<1M0 likes428 downloads11d agoHugging Face05ymoslem /wmt-da-human-evaluation-long-context Dataset Summary Long-context / document-level dataset for Quality Estimation of Machine Translation. It is an augmented variant of the sentence-level WMT DA Human Evaluation dataset. In addition to individual sentences, it contains augmentations of 2, 4, 8, 16, and 32 sentences, among each language pair lp and domain. The raw column represents a weighted average of scores of augmented sentences using character lengths of src and mt as weights. The code used to apply the augmentation… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/wmt-da-human-evaluation-long-context.tabular1M<n<10M8 likes398 downloads2y agoHugging Face06placeholderlabs /pretrain-commits-v2-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 12,571,681,749 (12.6B) Trainable tokens 4,460,160,435 (4.5B) Documents 992,475 Shards 327 UTF-8 bytes 49,288,867,997 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-commits-v2-mix-long-context.tabular1M<n<10M0 likes348 downloads6d agoHugging Face07placeholderlabs /pretrain-ultra-fineweb-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 1,367,358,024 (1.4B) Trainable tokens 1,367,358,024 (1.4B) Documents 48,077 Shards 73 UTF-8 bytes 6,386,740,105 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-ultra-fineweb-mix-long-context.tabular10K<n<100K1 likes322 downloads11d agoHugging Face08placeholderlabs /pretrain-academic-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 43,694,042,993 (43.7B) Trainable tokens 43,694,042,993 (43.7B) Documents 1,001,557 Shards 373 UTF-8 bytes 183,279,720,921 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-academic-mix-long-context.tabular1M<n<10M0 likes160 downloads11d agoHugging Face09Ashima /qwen3_0.6-task738_augmented_long-context_Mar16-1501tabular1K<n<10K0 likes117 downloads6mo agoHugging Face10placeholderlabs /pretrain-encyclopedic-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 587,625,128 (587.6M) Trainable tokens 587,625,128 (587.6M) Documents 23,631 Shards 9 UTF-8 bytes 1,978,753,989 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-encyclopedic-mix-long-context.tabular10K<n<100K1 likes101 downloads11d agoHugging Face11jannalu /mbpp-longcontext MBPP Long-Context Dataset Overview MBPP Long-Context is a benchmark dataset that combines coding problems from the MBPP (Mostly Basic Python Problems) dataset with long-context distractors from BABILong. This dataset evaluates code generation performance under long-context conditions, testing whether models can maintain coding ability with stuffed context. Dataset Structure Data Fields Each sample contains: Original MBPP Fields… See the full description on the dataset page: https://huggingface.co/datasets/jannalu/mbpp-longcontext.tabulartext-generation10K<n<100K0 likes90 downloads11mo agoHugging Face12placeholderlabs /pretrain-repository-v2-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 130,158,375,824 (130.2B) Trainable tokens 130,158,375,824 (130.2B) Documents 2,578,578 Shards 1,168 UTF-8 bytes 535,260,344,241 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-repository-v2-mix-long-context.tabular1M<n<10M0 likes89 downloads6d agoHugging Face13VivekChauhan06 /repobench_python_long_context RepoBench v1.1 (Python) Introduction This dataset presents the Python portion of RepoBench v1.1 (ICLR 2024). The data encompasses a collection from GitHub, spanning the period from October 6th to December 31st, 2023. With a commitment to data integrity, a deduplication process based on file content has been implemented against the Stack v2 dataset (coming soon), aiming to mitigate data leakage and memorization concerns. Features The dataset provides the… See the full description on the dataset page: https://huggingface.co/datasets/VivekChauhan06/repobench_python_long_context.tabular10K<n<100K1 likes70 downloads2y agoHugging Face14placeholderlabs /pretrain-nemotron-math-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 1,446,296,439 (1.4B) Trainable tokens 1,446,296,439 (1.4B) Documents 42,379 Shards 23 UTF-8 bytes 4,965,563,314 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-nemotron-math-mix-long-context.tabular10K<n<100K0 likes64 downloads11d agoHugging Face15aixsatoshi /Longcontext-aozora-instruction長文用のinstructionデータセットです。 長文は以下の青空文庫データセットを利用しました。 globis-university/aozorabunko-clean Limitation このデータセットは、長文の質問応答スタイルを提示することを主な目的としています。質問応答の正誤についてのフィルタリングはあえて行っていません。 長文では一般に性能低下が認められるため困難なタスクとなります。フィルタリングすると困難なタスクのinstructionが消えてしまうためです。ファインチューニングで使用する場合は、チューニングする基盤モデルの性能によって、チューニング効果が大きく変わります。正答できるかどうかはモデルパラメータ、事前学習次第と考えられます。 License CC BY 4.0 tabular1K<n<10K8 likes54 downloads2y agoHugging Face16Ashima /qwen3_0.6b_long-context_Mar17-1945_blendedtabularn<1K0 likes47 downloads6mo agoHugging Face17fineset-io /long-context-llm-papers Long-Context LLM Papers — FineSet A research-paper dataset on Long-Context LLM Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on Long-Context LLM Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset Quality-scored: quality_score… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/long-context-llm-papers.tabulartext-classificationn<1K0 likes36 downloads3mo agoHugging Face18Ashima /qwen3_0.6-task738_augmented_long-context_Mar16-1501_blendedtabularn<1K0 likes33 downloads6mo agoHugging Face19pritamdeb68 /FineWeb-long-context-documentstabular100K<n<1M1 likes29 downloads1y agoHugging Face20Kyle1668 /sfm-midtraining-mix-dclm-long-context-passages-blocklist-filteredtabular10K<n<100K0 likes29 downloads10mo agoHugging Face21NewEden /DCLM-Long-Context-Subsettabular10K<n<100K0 likes23 downloads10mo agoHugging Face22davanstrien /dataset_cards_with_long_context_embeddins Dataset Card for "dataset_cards_with_long_context_embeddins" More Information needed tabular10K<n<100K0 likes22 downloads3y agoHugging Face23tilde-research /long-contexttabularn<1K1 likes21 downloads1y agoHugging Face24davanstrien /model_cards_with_long_context_embeddings Dataset Card for "model_cards_with_long_context_embeddings" More Information needed tabular10K<n<100K0 likes18 downloads3y agoHugging Face25TAUR-dev /D-ExpTracker__1022_longcontext__maxlen4096_0epoch_3and4arg__v1tabularn<1K0 likes18 downloads11mo agoHugging Face26nreHieW /BigCodeBench-corrupted-long-context-no-teststabularn<1K0 likes17 downloads5mo agoHugging Face27GeneralAnalysis /GA_Long_Context_Jailbreak_Benchmarkgated GA Long Context Bench A benchmark of 1500 multi-turn conversations designed to stress guardrails in long contexts. Each dialog pairs a serialized agent trace with optional prompt injections or per-policy adjudications. Half of the rows embed malicious content deep inside long instructions, enabling evaluation of long-context systems. Accompanying guardrail releases: GA Guard Core and GA Guard Lite. Check out public benchmarks and results in our blogpost. [!Note] Disclaimer: This… See the full description on the dataset page: https://huggingface.co/datasets/GeneralAnalysis/GA_Long_Context_Jailbreak_Benchmark.tabular1K<n<10K2 likes17 downloads1y agoHugging Face28ttn0011 /longcontext_cottabularn<1K0 likes15 downloads1y agoHugging Face29TAUR-dev /D-ExpTracker__1022_longcontext__maxlen8192_1e_3args__v1tabularn<1K0 likes14 downloads11mo agoHugging Face30TAUR-dev /D-ExpTracker__1022_longcontext__maxlen4096_0epoch_3args__v1tabularn<1K0 likes13 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.