CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01huggingface /documentation-images This dataset contains images used in the documentation of HuggingFace's libraries. HF Team: Please make sure you optimize the assets before uploading them. My favorite tool for this is https://tinypng.com/. imagen<1K204 likes2m downloads8h agoHugging Face02hf-doc-build /doc-build-devThis is a dataset which contains the docs from all the PRs that are updating one of the docs from https://huggingface.co/docs. It is automatically updated by this github action from the doc-buider repo. 59 likes613k downloads5mo agoHugging Face03huggingface-course /documentation-imagesimagen<1K3 likes306k downloads1y agoHugging Face04hf-doc-build /doc-buildThis repo contains all the docs published on https://huggingface.co/docs. The docs are generated with https://github.com/huggingface/doc-builder. 41 likes228k downloads3h agoHugging Face05AmazonScience /document-haystack Document Haystack Dataset This repository contains the dataset for the paper “Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark”. 📑 Abstract Paper The proliferation of multimodal Large Language Models has significantly advanced the ability to analyze and understand complex data inputs from different modalities. However, the processing of long documents remains under-explored, largely due to a lack of suitable benchmarks. To… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/document-haystack.textquestion-answering20 likes89k downloads1y agoHugging Face06EleutherAI /wikitext_document_level Wikitext Document Level This is a modified version of https://huggingface.co/datasets/wikitext that returns Wiki pages instead of Wiki text line-by-line. The original readme is contained below. Dataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/wikitext_document_level.text10K<n<100K18 likes76k downloads2y agoHugging Face07racineai /VDR_MEGA_MultiDomain_DocRetrieval Visual Document Retrieval Dataset Overview This dataset is designed for training visual document retrieval models. It combines multiple datasets from the VDR series, Colpali, and LlamaIndex to create the most comprehensive training resource for visual document retrieval tasks. Dataset Structure The dataset contains structured fields including unique identifiers with string lengths ranging from 45 to 50 characters, search query text with variable lengths between… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_MultiDomain_DocRetrieval.imagevisual-document-retrieval1M<n<10M24 likes72k downloads6mo agoHugging Face08Xenova /transformers.js-docs3 likes71k downloads6mo agoHugging Face09HuggingFaceM4 /Docmatix Dataset Card for Docmatix Dataset description Docmatix is part of the Idefics3 release (stay tuned). It is a massive dataset for Document Visual Question Answering that was used for the fine-tuning of the vision-language model Idefics3. Load the dataset To load the dataset, install the library datasets with pip install datasets. Then, from datasets import load_dataset ds = load_dataset("HuggingFaceM4/Docmatix") If you want the dataset to link to the pdf files… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/Docmatix.imagevisual-question-answering1M<n<10M312 likes48k downloads2y agoHugging Face10trl-lib /documentation-imagesimagen<1K0 likes47k downloads19d agoHugging Face11lmms-lab-encoder /DocVQA Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of DocVQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @article{mathew2020docvqa, title={DocVQA: A Dataset for VQA on Document Images. CoRR abs/2007.00398 (2020)}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/DocVQA.image10K<n<100K87 likes37k downloads2y agoHugging Face12pruna-test /documentation-mediaimagen<1K0 likes35k downloads7d agoHugging Face13atom-in-the-universe /cc-doc-links0 likes35k downloads3y agoHugging Face14Felix92 /docTR-resource-collectiontext1K<n<10K1 likes34k downloads8mo agoHugging Face15jinaai /documentation-imagesimagen<1K0 likes28k downloads1y agoHugging Face16HPLT /DocHPLT DocHPLT: A Massively Multilingual Document-Level Translation Dataset Existing document-level machine translation resources are only available for a handful of languages, mostly high-resourced ones. To facilitate the training and evaluation of document-level translation and, more broadly, long-context modeling for global communities, we create DocHPLT, the largest publicly available document-level translation dataset to date. It contains 124 million aligned document pairs across 50… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/DocHPLT.texttranslation100M<n<1B20 likes15k downloads8mo agoHugging Face17huggingface /policy-docs Public Policy at Hugging Face AI Policy at Hugging Face is a multidisciplinary and cross-organizational workstream. Instead of being part of a vertical communications or global affairs organization, our policy work is rooted in the expertise of our many researchers and developers, from Ethics and Society Regulars and legal team to machine learning engineers working on healthcare, art, and evaluations. What we work on is informed by our Hugging Face community needs and experiences… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/policy-docs.documentn<1K15 likes13k downloads6mo agoHugging Face18thunlp /docredMultiple entities in a document generally exhibit complex inter-sentence relations, and cannot be well handled by existing relation extraction (RE) methods that typically focus on extracting intra-sentence relations for single entity pairs. In order to accelerate the research on document-level RE, we introduce DocRED, a new dataset constructed from Wikipedia and Wikidata with three features: - DocRED annotates both named entities and relations, and is the largest human-annotated dataset for document-level RE from plain text. - DocRED requires reading multiple sentences in a document to extract entities and infer their relations by synthesizing all information of the document. - Along with the human-annotated data, we also offer large-scale distantly supervised data, which enables DocRED to be adopted for both supervised and weakly supervised scenarios.text-retrieval100K<n<1M26 likes11k downloads3y agoHugging Face19diffusers /docs-imagesimagen<1K0 likes10k downloads6mo agoHugging Face20yubo2333 /MMLongBench-Docdocument1K<n<10K27 likes9.5k downloads11mo agoHugging Face21Marti844 /SaaS-Bench-docker SaaS-Bench Docker Images Docker image archives for the SaaS-Bench benchmark — a suite of 23 self-hosted SaaS applications used to evaluate computer-use LLM agents on real, multi-step business workflows. This repository hosts the prebuilt .tar images (≈ 63 GB total) so you can reproduce the benchmark environment without rebuilding each app from source. The eval harness, task definitions, and verifiers live in the main SaaS-Bench repository. Paper: SaaS-Bench: Can Computer-Use… See the full description on the dataset page: https://huggingface.co/datasets/Marti844/SaaS-Bench-docker.othern<1K0 likes9.3k downloads6d agoHugging Face22diffusers /diffusers-images-docsimagen<1K0 likes7.8k downloads2y agoHugging Face23tiiuae /documentation-imagesimagen<1K0 likes7.4k downloads1y agoHugging Face24docling-project /regression-dataset-for-docling-parse Regression Dataset for docling-parse This repository contains the reference dataset used as a regression test corpus for docling-parse. Its purpose is to make parser and renderer changes safe: when behavior changes in docling-parse, the test suite can compare the current output against the expected artifacts stored in this dataset. Correct workflow to add new files cp /path/to/new.pdf regression/new.pdf git add regression/new.pdf git lfs status git commit -s -m… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/regression-dataset-for-docling-parse.documentn<1K2 likes7.1k downloads7h agoHugging Face25optimum /documentation-imagesThis dataset contains images used in the documentation of HuggingFace's Optimum library. imagen<1K2 likes6.4k downloads2mo agoHugging Face26abigailhaddad /foia-reading-room-documents Foia Reading Room Documents Documents from federal FOIA reading rooms and Inspector General report libraries: audits, inspections, investigative summaries and records released under the Freedom of Information Act. Every document here was published by a US federal agency and is a work of the United States government. Nothing has been altered: files are byte-identical to what the agency posted, and the checksum in metadata.parquet is of the original bytes. Why this… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/foia-reading-room-documents.text-retrieval0 likes5.8k downloads9d agoHugging Face27HuggingFaceM4 /DoclingMatix DoclingMatix DoclingMatix is a large-scale, multimodal dataset designed for training vision-language models in the domain of document intelligence. It was created specifically for training the SmolDocling model, an ultra-compact model for end-to-end document conversion. The dataset is constructed by augmenting Hugging Face's Docmatix. Each sample in Docmatix, which consists of a document image and a few questions and answers about it, has been transformed. The text field is now… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/DoclingMatix.imagevisual-question-answering1M<n<10M56 likes5.1k downloads1y agoHugging Face28HCAI-Lab-GT /dolma3-6t-sample-100000-docs dolma3-6t-sample-100000-docs Materialized stratified sample of 100,000 docs per bin (~58k source-shard files, ~186 GB) from the deduplicated 6T Dolma3 corpus. Seed 42, materialized by 128 Modal workers (worker_0000/ through worker_0127/). Layout HCAI-Lab/dolma3-6t-sample-100000-docs/ ├── bin_summary.csv ├── sample_contract.json └── worker_NNNN/ └── soc127__phase1_pool_shared__...__shard_NNNNNNNN.jsonl.zst (~450 files per worker) Dual-access… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-100000-docs.0 likes5.1k downloads4mo agoHugging Face29amitbcp /docinsights-2026-shared-task-data DocInsights 2026 Shared Task: DocSem Document-grounded quantitative reasoning with evidence attribution DocSem is the shared task of DocInsights 2026, the Workshop on Document Intelligence and Understanding co-located with EMNLP 2026 in Budapest, Hungary. The workshop theme is Beyond Plain Text: Bridging NLP and Document AI. Workshop shared task | Source repository | Submission portal | Participant guide Participants receive a PDF document and a paraphrased user_query. Systems… See the full description on the dataset page: https://huggingface.co/datasets/amitbcp/docinsights-2026-shared-task-data.documentquestion-answering1K<n<10K0 likes4.9k downloads16d agoHugging Face30docling-project /DocLayNet-v1.2 Dataset Card for DocLayNet v1.2 Dataset Summary This dataset is an extention of the original DocLayNet dataset which embeds the PDF files of the document images inside a binary column. DocLayNet provides page-by-page layout segmentation ground-truth using bounding-boxes for 11 distinct class labels on 80863 unique pages from 6 document categories. It provides several unique features compared to related work such as PubLayNet or DocBank: Human Annotation: DocLayNet is… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/DocLayNet-v1.2.image10K<n<100K19 likes4.8k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.