CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amitbcp /docinsights-2026-shared-task-data DocInsights 2026 Shared Task: DocSem Document-grounded quantitative reasoning with evidence attribution DocSem is the shared task of DocInsights 2026, the Workshop on Document Intelligence and Understanding co-located with EMNLP 2026 in Budapest, Hungary. The workshop theme is Beyond Plain Text: Bridging NLP and Document AI. Workshop shared task | Source repository | Submission portal | Participant guide Participants receive a PDF document and a paraphrased user_query. Systems… See the full description on the dataset page: https://huggingface.co/datasets/amitbcp/docinsights-2026-shared-task-data.documentquestion-answering1K<n<10K0 likes4.9k downloads18d agoHugging Face02mannycooper /document-review-data Document Review Data Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package. Current Title Extraction Dataset Surface Canonical prefix: datasets/title_extraction/ Effective datasets: datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/ datasets/title_extraction/evaluation/real_device_280_v1/ datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/ The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.tabular100K<n<1M3 likes4.4k downloads13d agoHugging Face03Tongyi-Zhiwen /docmathtextn<1K0 likes3k downloads1y agoHugging Face04datasets-examples /doc-formats-jsonl-1 [doc] formats - jsonl - 1 This dataset contains one jsonl file at the root. textn<1K0 likes2.5k downloads3y agoHugging Face05mPLUG /DocDownstream-2.0DocDownstream 2.0 is a collection of MP-DocVQA, DUDE, NewsVideoQA used in DocOwl2. For MP-DocVQA and DUDE, the maximum number of pages of each sample is set to 20. For NewsVideoQA, the maximum number of frames of each sample is set to 20. image10K<n<100K3 likes1.9k downloads2y agoHugging Face06microsoft /XL-DocBench XL-DocBench Evidence-grounded reasoning across hundreds or thousands of pages. Fully verified by 194 human experts. Hongchen Wei1,†,‡, Yuanzhe Wang2,†,‡, Bei Liu2,*, Yifan Yang2, Qi Dai2, Ruichun Ma2, Kai Qiu2, Yunsheng Li2, Dongdong Chen2, Chong Luo2, Zhenzhong Chen1, Baining Guo2 1Wuhan University &nbsp; 2Microsoft &nbsp; †Equal contribution &nbsp; ‡Work done during an internship at MSRA &nbsp; *Project leader Project Page · Paper · Live Leaderboard… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/XL-DocBench.tabularquestion-answering1K<n<10K7 likes960 downloads22d agoHugging Face07kensho /DocFinQAtext1K<n<10K21 likes925 downloads2y agoHugging Face08philschmid /markdown-documentation-transformers Hugging Face Transformers documentation as markdown dataset This dataset was created using Clipper.js. Clipper is a Node.js command line tool that allows you to easily clip content from web pages and convert it to Markdown. It uses Mozilla's Readability library and Turndown under the hood to parse web page content and convert it to Markdown. This dataset can be used to create RAG applications, which want to use the transformers documentation. Example document:… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/markdown-documentation-transformers.textn<1K11 likes922 downloads3y agoHugging Face09nyuuzyou /znanio-documents Dataset Card for Znanio.ru Educational Documents Dataset Summary This dataset contains 588,545 educational document files from the znanio.ru platform, a resource for teachers, educators, students, and parents providing diverse educational content. Znanio.ru has been a pioneer in educational technologies and distance learning in the Russian-speaking internet since 2009. The dataset includes a small portion of English language content, primarily for language learning… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/znanio-documents.texttext-classification100K<n<1M0 likes902 downloads2y agoHugging Face10tonytan48 /Re-DocRED Re-DocRED Dataset This repository contains the dataset of our EMNLP 2022 research paper Revisiting DocRED – Addressing the False Negative Problem in Relation Extraction. DocRED is a widely used benchmark for document-level relation extraction. However, the DocRED dataset contains a significant percentage of false negative examples (incomplete annotation). We revised 4,053 documents in the DocRED dataset and resolved its problems. We released this dataset as: Re-DocRED dataset. The… See the full description on the dataset page: https://huggingface.co/datasets/tonytan48/Re-DocRED.text1K<n<10K5 likes760 downloads4y agoHugging Face11AIR-Bench /long-doc_book_enAvailable Versions: AIR-Bench_24.04 Task / Domain / Language: long-doc / book / en Available Datasets (Dataset Name: Splits): origin-of-species_darwin: test a-brief-history-of-time_stephen-hawking: test AIR-Bench_24.05 Task / Domain / Language: long-doc / book / en Available Datasets (Dataset Name: Splits): origin-of-species_darwin: test a-brief-history-of-time_stephen-hawking: dev texttext-retrieval1K<n<10K0 likes635 downloads2y agoHugging Face12docketx /us-caselaw US Caselaw Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them. The American case-law record in one dataset — published opinions from the Free Law Project / CourtListener bulk export of 2026-06-30: full text, all opinion types (majority, concurring, dissenting, per curiam… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-caselaw.texttext-retrieval1M<n<10M0 likes616 downloads3d agoHugging Face13r2e-edits /dockersv1tabular1K<n<10K0 likes578 downloads2y agoHugging Face14semeru /text-code-galeras-code-generation-from-docstring-3k-dedupedtabular1K<n<10K0 likes577 downloads3y agoHugging Face15AmazonScience /DocTalk 📁 DocTalk: Scalable Graph-based Dialogue Synthesis for Enhancing LLM Conversational Capabilities ➤ 📖 Paper Link DocTalk is a large-scale, synthetic dialogue corpus created via a three-stage pipeline that converts clusters of related Wikipedia documents into multi-turn, multi-topic information-seeking conversations. The pipeline comprises: Document Graph Construction: Sampling up to three related Wikipedia articles per anchor document via a weighted random walk on a directed… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/DocTalk.textquestion-answering100K<n<1M2 likes549 downloads1y agoHugging Face16rmems /docker-build-cache-trajectories Docker Build Cache Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/docker-build-cache-trajectories.text1K<n<10K0 likes523 downloads16h agoHugging Face17gradio /docstext1K<n<10K3 likes493 downloads20h agoHugging Face18AquaV /mil-docs What is this? A curated selection of manuals and documents from the US military and other departments. All data was manually scraped from publicly available sources. The PDF's and EPUB files were converted to markdown using the amazing Marker github repository by Vik Paruchuri. Sources: United States Army Central Army Repository Marines Publications Federation of American Scientists Intelligence Resource Program text1K<n<10K2 likes483 downloads3y agoHugging Face19agentlans /en-document-classification English Document Classification Dataset This dataset provides a curated subset of the first 1 million rows from the allenai/c4 (English configuration), enriched with multi-perspective topic annotations. It is designed for researchers exploring document classification, domain adaptation, and label noise in massive web-crawled corpora. Dataset Summary The dataset integrates predictions from classification models to provide a holistic view of each document’s content.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-classification.texttext-classification1M<n<10M1 likes419 downloads24d agoHugging Face20mannycooper /document-review-source1k New 1K title extraction corpus Independent Task0178 expansion, parsed and machine annotated in Task0180, human-reviewed under Task0186. This is NOT the 1K subset sampled from the older 4K corpus. Private research use only; original document rights are not independently verified. 1000 unique source documents: Word, Excel, PDF and PPT each 250; English 500, simplified Chinese 400, traditional Chinese 100. All 1000 human reviews collected: 711 single-title, 194 multi-title, 54… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-source1k.document1K<n<10K0 likes415 downloads1d agoHugging Face21hazyresearch /LoCoV1-Documents LoCoV1 Documents The documents for the LoCoV1 dataset from the paper, "Benchmarking and Building Long-Context Retrieval Models with LoCo and M2-BERT" How to Use To load the dataset, use the following command: from datasets import load_dataset dataset = load_dataset("hazyresearch/LoCoV1-Documents") To load a specific subset, such as SummScreenFD, use the following command: from datasets import load_dataset dataset = load_dataset("hazyresearch/LoCoV1-Documents") def… See the full description on the dataset page: https://huggingface.co/datasets/hazyresearch/LoCoV1-Documents.text10K<n<100K5 likes409 downloads3y agoHugging Face22docketx /texas-caselaw Texas Case Law Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them. Source & credit — Free Law Project / CourtListener Every opinion in this dataset comes from the Free Law Project / CourtListener bulk export of 2026-06-30 (10,798,347 opinions). CourtListener… See the full description on the dataset page: https://huggingface.co/datasets/docketx/texas-caselaw.texttext-retrieval100K<n<1M0 likes367 downloads3d agoHugging Face23pgurazada1 /document-qna-chroma-anyscale-logstextn<1K0 likes363 downloads2y agoHugging Face24docketx /us-regulations US Federal Regulations — the CFR, held word for word Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them. 222,767 regulations — 219,061 CFR sections and 3,706 appendices — across all 49 titles of the Code of Federal Regulations, each one as the agency publishes it. Statutes say… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-regulations.texttext-retrieval100K<n<1M0 likes344 downloads3d agoHugging Face25docketx /us-caselaw-ca California Case Law Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them. Full text of 242,989 California appellate opinion documents from the public record. The base slice is 242,831 documents from the Free Law Project / CourtListener bulk export of 2026-06-30.… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-caselaw-ca.texttext-retrieval100K<n<1M0 likes314 downloads3d agoHugging Face26glaiveai /godot_4_docsDataset generated for Godot 4 docs using Glaive. text1K<n<10K21 likes309 downloads2y agoHugging Face27docketx /us-statutes US Statutes — held word for word Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them. 1,295,620 statute sections: the whole United States Code (60,433 sections, all 53 titles) plus 27 states (1,235,187 sections), each section as its legislature publishes it, with the URL it was… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-statutes.texttext-retrieval1M<n<10M0 likes309 downloads3d agoHugging Face28saidsef /tech-docs Technical Documentation Dataset A curated collection of technical documentation and guides spanning various cloud-native technologies, infrastructure tools, and machine learning frameworks. This dataset contains 1,397 documents in JSONL format, covering essential topics for modern software development and DevOps practices. Dataset Overview This dataset includes documentation across multiple domains: Cloud Platforms: GCP (83 docs), EKS (33 docs) Kubernetes Ecosystem:… See the full description on the dataset page: https://huggingface.co/datasets/saidsef/tech-docs.textquestion-answering1K<n<10K2 likes308 downloads2y agoHugging Face29joshycodes /sorrel-T-qwen3-8b-base-seed0-documentstext100K<n<1M0 likes304 downloads5d agoHugging Face30docketx /us-caselaw-wi Wisconsin Case Law Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them. Full text of 72,350 Wisconsin appellate opinion documents from the public record, sliced from the Free Law Project / CourtListener bulk export of 2026-06-30. Court coverage (2 court ids, explicit allowlist —… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-caselaw-wi.texttext-retrieval10K<n<100K0 likes303 downloads3d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.