CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amitbcp /docinsights-2026-shared-task-data DocInsights 2026 Shared Task: DocSem Document-grounded quantitative reasoning with evidence attribution DocSem is the shared task of DocInsights 2026, the Workshop on Document Intelligence and Understanding co-located with EMNLP 2026 in Budapest, Hungary. The workshop theme is Beyond Plain Text: Bridging NLP and Document AI. Workshop shared task | Source repository | Submission portal | Participant guide Participants receive a PDF document and a paraphrased user_query. Systems… See the full description on the dataset page: https://huggingface.co/datasets/amitbcp/docinsights-2026-shared-task-data.documentquestion-answering1K<n<10K0 likes4.9k downloads16d agoHugging Face02mannycooper /document-review-data Document Review Data Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package. Current Title Extraction Dataset Surface Canonical prefix: datasets/title_extraction/ Effective datasets: datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/ datasets/title_extraction/evaluation/real_device_280_v1/ datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/ The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.tabular100K<n<1M3 likes3k downloads11d agoHugging Face03datasets-examples /doc-formats-jsonl-1 [doc] formats - jsonl - 1 This dataset contains one jsonl file at the root. textn<1K0 likes2.4k downloads3y agoHugging Face04Tongyi-Zhiwen /docmathtextn<1K0 likes1.3k downloads1y agoHugging Face05philschmid /markdown-documentation-transformers Hugging Face Transformers documentation as markdown dataset This dataset was created using Clipper.js. Clipper is a Node.js command line tool that allows you to easily clip content from web pages and convert it to Markdown. It uses Mozilla's Readability library and Turndown under the hood to parse web page content and convert it to Markdown. This dataset can be used to create RAG applications, which want to use the transformers documentation. Example document:… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/markdown-documentation-transformers.textn<1K11 likes942 downloads3y agoHugging Face06microsoft /XL-DocBench XL-DocBench Evidence-grounded reasoning across hundreds or thousands of pages. Fully verified by 194 human experts. Hongchen Wei1,†,‡, Yuanzhe Wang2,†,‡, Bei Liu2,*, Yifan Yang2, Qi Dai2, Ruichun Ma2, Kai Qiu2, Yunsheng Li2, Dongdong Chen2, Chong Luo2, Zhenzhong Chen1, Baining Guo2 1Wuhan University &nbsp; 2Microsoft &nbsp; †Equal contribution &nbsp; ‡Work done during an internship at MSRA &nbsp; *Project leader Project Page · Paper · Live Leaderboard… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/XL-DocBench.tabularquestion-answering1K<n<10K7 likes932 downloads20d agoHugging Face07nyuuzyou /znanio-documents Dataset Card for Znanio.ru Educational Documents Dataset Summary This dataset contains 588,545 educational document files from the znanio.ru platform, a resource for teachers, educators, students, and parents providing diverse educational content. Znanio.ru has been a pioneer in educational technologies and distance learning in the Russian-speaking internet since 2009. The dataset includes a small portion of English language content, primarily for language learning… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/znanio-documents.texttext-classification100K<n<1M0 likes889 downloads2y agoHugging Face08kensho /DocFinQAtext1K<n<10K21 likes742 downloads2y agoHugging Face09tonytan48 /Re-DocRED Re-DocRED Dataset This repository contains the dataset of our EMNLP 2022 research paper Revisiting DocRED – Addressing the False Negative Problem in Relation Extraction. DocRED is a widely used benchmark for document-level relation extraction. However, the DocRED dataset contains a significant percentage of false negative examples (incomplete annotation). We revised 4,053 documents in the DocRED dataset and resolved its problems. We released this dataset as: Re-DocRED dataset. The… See the full description on the dataset page: https://huggingface.co/datasets/tonytan48/Re-DocRED.text1K<n<10K5 likes730 downloads4y agoHugging Face10AIR-Bench /long-doc_book_enAvailable Versions: AIR-Bench_24.04 Task / Domain / Language: long-doc / book / en Available Datasets (Dataset Name: Splits): origin-of-species_darwin: test a-brief-history-of-time_stephen-hawking: test AIR-Bench_24.05 Task / Domain / Language: long-doc / book / en Available Datasets (Dataset Name: Splits): origin-of-species_darwin: test a-brief-history-of-time_stephen-hawking: dev texttext-retrieval1K<n<10K0 likes609 downloads2y agoHugging Face11docketx /us-caselaw US Caselaw Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them. The American case-law record in one dataset — published opinions from the Free Law Project / CourtListener bulk export of 2026-06-30: full text, all opinion types (majority, concurring, dissenting, per curiam… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-caselaw.texttext-retrieval1M<n<10M0 likes594 downloads19h agoHugging Face12r2e-edits /dockersv1tabular1K<n<10K0 likes573 downloads2y agoHugging Face13semeru /text-code-galeras-code-generation-from-docstring-3k-dedupedtabular1K<n<10K0 likes565 downloads3y agoHugging Face14AmazonScience /DocTalk 📁 DocTalk: Scalable Graph-based Dialogue Synthesis for Enhancing LLM Conversational Capabilities ➤ 📖 Paper Link DocTalk is a large-scale, synthetic dialogue corpus created via a three-stage pipeline that converts clusters of related Wikipedia documents into multi-turn, multi-topic information-seeking conversations. The pipeline comprises: Document Graph Construction: Sampling up to three related Wikipedia articles per anchor document via a weighted random walk on a directed… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/DocTalk.textquestion-answering100K<n<1M2 likes517 downloads1y agoHugging Face15rmems /docker-build-cache-trajectories Docker Build Cache Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/docker-build-cache-trajectories.text1K<n<10K0 likes497 downloads21d agoHugging Face16gradio /docstext1K<n<10K3 likes488 downloads3d agoHugging Face17AquaV /mil-docs What is this? A curated selection of manuals and documents from the US military and other departments. All data was manually scraped from publicly available sources. The PDF's and EPUB files were converted to markdown using the amazing Marker github repository by Vik Paruchuri. Sources: United States Army Central Army Repository Marines Publications Federation of American Scientists Intelligence Resource Program text1K<n<10K2 likes483 downloads3y agoHugging Face18agentlans /en-document-classification English Document Classification Dataset This dataset provides a curated subset of the first 1 million rows from the allenai/c4 (English configuration), enriched with multi-perspective topic annotations. It is designed for researchers exploring document classification, domain adaptation, and label noise in massive web-crawled corpora. Dataset Summary The dataset integrates predictions from classification models to provide a holistic view of each document’s content.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-classification.texttext-classification1M<n<10M1 likes426 downloads22d agoHugging Face19Doc2Feat-bench /Doc2Feat-bench_Verified Dataset Summary NoCode-bench Verified is subset of NoCode-bench, a dataset that tests systems’ no-code feature addition ability automatically. Languages The text of the dataset is primarily English, but we make no effort to filter or otherwise clean based on language type. Dataset Structure An example of a SWE-bench datum is as follows: repo: (str) - The repository owner/name identifier from GitHub. instance_id: (str) - A formatted instance… See the full description on the dataset page: https://huggingface.co/datasets/Doc2Feat-bench/Doc2Feat-bench_Verified.texttext-generationn<1K1 likes420 downloads1y agoHugging Face20hazyresearch /LoCoV1-Documents LoCoV1 Documents The documents for the LoCoV1 dataset from the paper, "Benchmarking and Building Long-Context Retrieval Models with LoCo and M2-BERT" How to Use To load the dataset, use the following command: from datasets import load_dataset dataset = load_dataset("hazyresearch/LoCoV1-Documents") To load a specific subset, such as SummScreenFD, use the following command: from datasets import load_dataset dataset = load_dataset("hazyresearch/LoCoV1-Documents") def… See the full description on the dataset page: https://huggingface.co/datasets/hazyresearch/LoCoV1-Documents.text10K<n<100K5 likes400 downloads2y agoHugging Face21pgurazada1 /document-qna-chroma-anyscale-logstextn<1K0 likes363 downloads2y agoHugging Face22saidsef /tech-docs Technical Documentation Dataset A curated collection of technical documentation and guides spanning various cloud-native technologies, infrastructure tools, and machine learning frameworks. This dataset contains 1,397 documents in JSONL format, covering essential topics for modern software development and DevOps practices. Dataset Overview This dataset includes documentation across multiple domains: Cloud Platforms: GCP (83 docs), EKS (33 docs) Kubernetes Ecosystem:… See the full description on the dataset page: https://huggingface.co/datasets/saidsef/tech-docs.textquestion-answering1K<n<10K2 likes334 downloads2y agoHugging Face23mPLUG /DocDownstream-2.0DocDownstream 2.0 is a collection of MP-DocVQA, DUDE, NewsVideoQA used in DocOwl2. For MP-DocVQA and DUDE, the maximum number of pages of each sample is set to 20. For NewsVideoQA, the maximum number of frames of each sample is set to 20. image10K<n<100K3 likes317 downloads2y agoHugging Face24docketx /texas-caselaw Texas Case Law Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them. Source & credit — Free Law Project / CourtListener Every opinion in this dataset comes from the Free Law Project / CourtListener bulk export of 2026-06-30 (10,798,347 opinions). CourtListener… See the full description on the dataset page: https://huggingface.co/datasets/docketx/texas-caselaw.texttext-retrieval100K<n<1M0 likes306 downloads19h agoHugging Face25joshycodes /sorrel-T-qwen3-8b-base-seed0-documentstext100K<n<1M0 likes301 downloads3d agoHugging Face26Anonymous-Team-HC-RAG /Multi-doc-2025 Dataset Card for Multi-Doc-2025 Dataset Summary Multi-Doc-2025 is a financial question-answering benchmark built from SEC Form 10-K annual reports of S&P 500 companies. It is designed to evaluate retrieval-augmented generation (RAG) and financial QA systems under three reasoning settings that are not jointly covered by existing financial QA benchmarks: cross-company reasoning, cross-year reasoning, and hybrid-modal reasoning over both text and tables. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-Team-HC-RAG/Multi-doc-2025.textquestion-answering1K<n<10K2 likes296 downloads4mo agoHugging Face27HCAI-Lab-GT /dolma3-6t-sample-10000-docs-finance-and-business HCAI-Lab/dolma3-6t-sample-10000-docs-finance-and-business Filename-derived finance_and_business slice of HCAI-Lab/dolma3-6t-sample-10000-docs, pinned to revision 561e73c7e0ad35c04f386bae1e3dd39dfb6755e7. Extraction rule The corpus contains every source .jsonl.zst file whose filename contains the literal segment -finance_and_business-. Source paths and compressed file contents are preserved byte-for-byte. This is a coarse WebOrganizer finance_and_business category… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-10000-docs-finance-and-business.texttext-generation100K<n<1M0 likes288 downloads1mo agoHugging Face28anonymous-8421 /VL-DocIRgated Abstract VL-DocIR is a page-level benchmark for vision-based long-document retrieval built from 29,641 documents rendered into 388,548 page images from Wikipedia, arXiv, PubMed, and SEC proxy statements. The benchmark contains 271,760 questions over 23 domains and six query types, covering single-page, multi-page, and cross-document evidence configurations. Questions are grounded to rendered pages and HTML element identifiers, then filtered with a cleaning pipeline that targets… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-8421/VL-DocIR.textvisual-document-retrieval1M<n<10M0 likes281 downloads5mo agoHugging Face29keisuke-miyako /doc-2026-0612text10K<n<100K0 likes281 downloads3mo agoHugging Face30joshycodes /sorrel-T-qwen3-1.7b-base-seed0-documentstext100K<n<1M0 likes280 downloads4d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.