CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mannycooper /document-review-data Document Review Data Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package. Current Title Extraction Dataset Surface Canonical prefix: datasets/title_extraction/ Effective datasets: datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/ datasets/title_extraction/evaluation/real_device_280_v1/ datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/ The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.tabular100K<n<1M3 likes3k downloads12d agoHugging Face02RoboCOIN /Agilex_Cobot_Magic_zip_up_the_document_baggated Agilex_Cobot_Magic_zip_up_the_document_bag 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: Agilex_Cobot_Magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: office 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pull place 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_zip_up_the_document_bag.tabularrobotics100K<n<1M0 likes1.2k downloads3mo agoHugging Face03singletongue /cc100-documents cc100-documents This dataset is a restructured version of the CC-100 (statmt/cc100) dataset. In the original dataset, each instance corresponds to a single paragraph (or a document boundary). In this version, the data has been reformed so that each instance corresponds to a single, complete document. This document-level structure makes it more convenient for processing using the map() and filter() methods in the Hugging Face Datasets library. Languages The following… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/cc100-documents.tabulartext-generation100M<n<1B1 likes1.1k downloads11mo agoHugging Face04RoboCOIN /RMC-AIDA-L_organise_the_document_baggated RMC-AIDA-L_organise_the_document_bag 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: realman_rmc_aidal | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: office 🤖 Atomic Actions This dataset includes the following atomic actions: grasp place pick pull 📊 Dataset Statistics Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/RMC-AIDA-L_organise_the_document_bag.tabularrobotics100K<n<1M0 likes939 downloads9mo agoHugging Face05jinaai /wikimedia-commons-documents-ml_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_beir.image1K<n<10K0 likes804 downloads1y agoHugging Face06DBWBD /Chinese_Debate_Documents Dataset Card for Chinese Debate Documents ASR-transcribed corpus of competitive Mandarin Chinese university debates, speaker-segmented and timestamped, with topic / round / team metadata parsed from the source filenames. Loading the Dataset from datasets import load_dataset ds = load_dataset("DBWBD/Chinese_Debate_Documents", split="train") print(ds[0]["topic"], "—", ds[0]["team_a"], "vs", ds[0]["team_b"]) for seg in ds[0]["segments"][:3]: print(f"… See the full description on the dataset page: https://huggingface.co/datasets/DBWBD/Chinese_Debate_Documents.tabulartext-classification1K<n<10K1 likes766 downloads4mo agoHugging Face07hulk10 /data_gouv_datasets_catalog-full-documents 🇫🇷 Catalogue des jeux de données de data.gouv.fr – Version structurée Ce dataset constitue une version structurée et exhaustive du catalogue des jeux de données publiés sur data.gouv.fr, la plateforme nationale française de l’open data. Il recense l’ensemble des jeux de données référencés sur la plateforme et fournit leurs métadonnées complètes : titre et description, organisation productrice, licence, couverture spatiale et temporelle, fréquence de mise à jour, formats… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/data_gouv_datasets_catalog-full-documents.tabular100K<n<1M0 likes526 downloads2h agoHugging Face08hulk10 /conseil-detat-full-documents Décisions du Conseil d'État (France) Description Ce dataset contient un corpus de décisions rendues par le Conseil d'État français, la plus haute juridiction de l'ordre administratif. Les décisions proviennent de la plateforme officielle Open Data de la Justice Administrative et sont diffusées au format XML anonymisé. Le corpus rassemble les textes intégraux des décisions ainsi que plusieurs métadonnées permettant leur identification et leur traçabilité. Source… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-detat-full-documents.tabularquestion-answering1K<n<10K1 likes513 downloads2h agoHugging Face09LegionIntel /named_entity_recognition_document_contexttabular1M<n<10M3 likes270 downloads2y agoHugging Face10dtunkelang /bag-of-documents Bag-of-Documents: Product Search Dataset Blog post: Distilling Retrieval Pipelines to a Single Embedding Model Live demo: huggingface.co/spaces/dtunkelang/bag-of-documents-demo Code: github.com/dtunkelang/bag-of-documents Dataset Description A large-scale bag-of-documents dataset for e-commerce product search, built on Amazon product data. Each search query is represented as a distribution of relevant products in embedding space, captured by a centroid vector and a… See the full description on the dataset page: https://huggingface.co/datasets/dtunkelang/bag-of-documents.tabularsentence-similarityn<1K5 likes250 downloads5mo agoHugging Face11histde /dta-documents Deutsches Textarchiv (DTA) Documents This datasets hosts all documents from the Deutsches Textarchiv (DTA). One row per work of the Deutsches Textarchiv (DTA), built from the official DTA dump dta_komplett_2026-02-10 (TCF format). The dataset contains 5,480 documents spanning 1472 to 1987 with about 205M tokens and 1.4B characters of German text, with two text views per document: text: the historical, layout-faithful transcription (line breaks, long s ſ, combining diacritics… See the full description on the dataset page: https://huggingface.co/datasets/histde/dta-documents.tabular1K<n<10K0 likes224 downloads1mo agoHugging Face12KanoonGPT /indian-legal-documents Indian Legal Documents Open Indian statutory and legal-document data for AI, search, and legal research. This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems. KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-legal-documents.texttext-generation10K<n<100K0 likes222 downloads4mo agoHugging Face13erdem-erdem /Turkish-Law-Documents-700k-clustered Turkish Legal Documents Clustering Dataset A comprehensive dataset of 700,000 Turkish legal documents from the two primary sources of legal precedent in Turkey, clustered using multiple emebdding models and algorithms to enable research, analysis, and machine learning applications. Overview This repository contains a large-scale document clustering pipeline and dataset for Turkish legal documents sourced from: Yargıtay - Turkey's highest court of appeal for civil and… See the full description on the dataset page: https://huggingface.co/datasets/erdem-erdem/Turkish-Law-Documents-700k-clustered.tabular100K<n<1M7 likes213 downloads11mo agoHugging Face14JBrightmanAI /TranNhiem-Vietnamese-DocumentImage-Reasoning TranNhiem Vietnamese Document-Image Reasoning (V-Doc) Vietnamese document-image understanding with explicit reasoning: multi-turn question–answering grounded on scanned/rendered Vietnamese document pages (textbooks, articles, worksheets). Each answer includes a step-by-step chain-of-thought. Reasoning and Answer was synthesized by Qwen3.5-397B-A17B over the Viet-Doc-VQA-II document collection. Curated by: Trần Nhiệm.. Mình rất welcome cho các hợp tác liên quan tới building… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/TranNhiem-Vietnamese-DocumentImage-Reasoning.imagevisual-question-answering10K<n<100K0 likes207 downloads2mo agoHugging Face15sheggle /dutch-legal-documents Dutch Legal Documents A comprehensive collection of 1.14 million Dutch legal documents, including court rulings, parliamentary documents, and EU legislation. Sources Source Type Documents Description Rechtspraak.nl Court rulings 882,212 All Dutch court rulings from Hoge Raad, Raad van State, Gerechtshoven, Rechtbanken, CRvB, CBb Officiële Bekendmakingen Parliamentary docs 255,837 Kamerstukken, Handelingen, Memories van Toelichting, Staatsblad EUR-Lex EU… See the full description on the dataset page: https://huggingface.co/datasets/sheggle/dutch-legal-documents.tabular1M<n<10M1 likes203 downloads6mo agoHugging Face16trannhiem /TranNhiem-Vietnamese-DocumentImage-Reasoning TranNhiem Vietnamese Document-Image Reasoning (V-Doc) Vietnamese document-image understanding with explicit reasoning: multi-turn question–answering grounded on scanned/rendered Vietnamese document pages (textbooks, articles, worksheets). Each answer includes a step-by-step chain-of-thought. Reasoning and Answer was synthesized by Qwen3.5 over the Viet-Doc-VQA-II document collection. Curated by: Trần Nhiệm.. Languages: Vietnamese (vi) answers · English (en) reasoning… See the full description on the dataset page: https://huggingface.co/datasets/trannhiem/TranNhiem-Vietnamese-DocumentImage-Reasoning.imagevisual-question-answering10K<n<100K5 likes173 downloads2mo agoHugging Face17ClimatePolicyRadar /all-document-text-datagated Climate Policy Radar Open Data This repo contains the full text data of all of the documents from the Climate Policy Radar database (CPR), which is also available at Climate Change Laws of the World (CCLW). Please note that this replaces the Global Stocktake open dataset: that data, including all NDCs and IPCC reports is now a subset of this dataset. What’s in this dataset This dataset contains two corpus types (groups of the same types or sources of documents) which… See the full description on the dataset page: https://huggingface.co/datasets/ClimatePolicyRadar/all-document-text-data.tabular10M<n<100M24 likes137 downloads11mo agoHugging Face18operant-ai /doclaynet-document-level DocLayNet Document-Level Reconstruction and 8K Expansion This dataset is a normalized, one-row-per-document view over the page-level DocLayNet v1.1 dataset. Pages are grouped using DocLayNet's source metadata and ordered by their original page number. Dataset summary 2,944 logical documents 80,863 observed pages 896 complete document groups 2,048 partial document groups Train: 2,355 documents / 60,810 pages Validation: 294 documents / 7,964 pages Test: 295… See the full description on the dataset page: https://huggingface.co/datasets/operant-ai/doclaynet-document-level.tabular10K<n<100K1 likes135 downloads11d agoHugging Face19MonumentalSystems /document-corpus-v3-open Document Corpus v3 Open document-corpus-v3-open is the redistribution-compatible slice of the exact byte-level pretraining corpus used by MonumentalSystems' 128M Harmonic GPT experiments. It contains 869,739 filtered documents and 2.192 GB of UTF-8 text before Parquet compression. This is not the complete internal document-corpus-v3. Restricted, unknown-license, and share-alike sources were excluded conservatively. Every included row comes from an upstream dataset whose card… See the full description on the dataset page: https://huggingface.co/datasets/MonumentalSystems/document-corpus-v3-open.tabulartext-generation100K<n<1M0 likes131 downloads2mo agoHugging Face20prg-unibe /dodis-historical-documents Dataset Overview The Swiss Historical Archive of DOdis Works (SHADOW) is a large-scale benchmark dataset derived from the resources of Dodis, an independent research center dedicated to the study of Swiss foreign policy and Switzerland’s international relations. Dodis has curated and published nearly 50,000 key documents that document the administrative practices and decision-making processes of the Swiss federal administration. The underlying corpus consists of notes, letters… See the full description on the dataset page: https://huggingface.co/datasets/prg-unibe/dodis-historical-documents.tabularsummarization10K<n<100K4 likes128 downloads3mo agoHugging Face21ClimatePolicyRadar /global-stocktake-documents Global Stocktake Open Data This repo contains the data for the first UNFCCC Global Stocktake. The data consists of document metadata from sources relevant to the Global Stocktake process, as well as full text parsed from the majority of the documents. The files in this dataset are as follows: metadata.csv: a CSV containing document metadata for each document we have collected. This metadata may not be the same as what's stored in the source databases – we have cleaned and added… See the full description on the dataset page: https://huggingface.co/datasets/ClimatePolicyRadar/global-stocktake-documents.tabular1M<n<10M7 likes120 downloads3y agoHugging Face22FudanSELab /SO_KGXQR_DOCUMENT Dataset Card for "SO_KGXQR_DOCUMENT" More Information needed tabular100K<n<1M0 likes118 downloads3y agoHugging Face23yourbench /aws_bedrock_documentation_demo Aws Bedrock Documentation Demo This dataset was generated using YourBench (v0.6.0), an open-source framework for generating domain-specific benchmarks from document collections. Pipeline Steps ingestion: Read raw source documents, convert them to normalized markdown and save for downstream steps summarization: Perform hierarchical summarization: chunk-level LLM summaries followed by combine-stage reduction chunking: Split texts into token-based single-hop and multi-hop… See the full description on the dataset page: https://huggingface.co/datasets/yourbench/aws_bedrock_documentation_demo.tabular1K<n<10K0 likes101 downloads1y agoHugging Face24PlethoraSolutions /open-scholarly-document-catalog-10k Search interactively &nbsp;·&nbsp; Product and methodology &nbsp;·&nbsp; Full catalog Rights-Audited Scholarly PDF Catalog: 10K Validated Sample Evaluation sample / inspect before you buy. Metadata and source links only. No PDFs or extracted full text are redistributed. Build scientific-literature search, RAG, discovery, and corpus-acquisition workflows from scholarly metadata with current PDF-link evidence and per-record Creative Commons or NASA rights… See the full description on the dataset page: https://huggingface.co/datasets/PlethoraSolutions/open-scholarly-document-catalog-10k.tabulartext-retrieval10K<n<100K0 likes90 downloads13d agoHugging Face25abby2231 /indian-legal-documents Indian Legal Documents Open Indian statutory and legal-document data for AI, search, and legal research. This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems. KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/abby2231/indian-legal-documents.texttext-generation10K<n<100K0 likes63 downloads1mo agoHugging Face26datamatters24 /research-document-archive Research Document Archive Computational analysis of 234,630 declassified U.S. government documents across 7 archival collections. Output of a 13-step ML pipeline extracting OCR text, entities, topics, keywords, redactions, and semantic embeddings from 3.1 million pages. Files File Rows Description documents.parquet 234,630 Document metadata: id, source_section, file_path, file_hash, total_pages pages/<section>.parquet 3.1M Per-page OCR text + 1536-dim… See the full description on the dataset page: https://huggingface.co/datasets/datamatters24/research-document-archive.tabulartext-classification10M<n<100M0 likes59 downloads5mo agoHugging Face27biglam /muninn-ww1-documents Muninn WWI Documents (CEF Attestation Papers & War Diaries) A tabular conversion of the document records in the Muninn Project's World War I Linked Open Data archive (https://rdf.muninn-project.org/). The Muninn Project is a research project that extracts structured data from digitized WWI-era archival documents. The dataset covers 28,742 documents — 28,718 Canadian Expeditionary Force (CEF) attestation papers (enlistment forms of individual soldiers) and 24 Canadian unit war… See the full description on the dataset page: https://huggingface.co/datasets/biglam/muninn-ww1-documents.tabularimage-classification100K<n<1M0 likes59 downloads2mo agoHugging Face28thoughtworks /document-processing-benchmark Document Processing Benchmark 8 public document datasets (receipts, invoices, forms, bank statements, multi-page docs, contracts) normalized into one parquet schema. Each row has the document, ground-truth annotations, and per-row token/latency/cost numbers from real API calls to one or more reference models. You can read off a target's cost/latency/quality without re-running it. from datasets import load_dataset ds = load_dataset("thoughtworks/document-processing-benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/document-processing-benchmark.tabularimage-to-text10K<n<100K1 likes58 downloads4mo agoHugging Face29timodonnell /bioreason-pro-sft-reasoning-documents BioReason-Pro SFT Reasoning Documents Complete, self-contained training documents reconstructed from the BioReason-Pro SFT data, ready for LLM pre-training. The upstream dataset wanglab/bioreason-pro-sft-reasoning-data ships the assistant side of each training example (reasoning, final_answer) alongside the raw biological context columns, but not the assembled prompt. The prompt cannot be recovered from the data card alone, because two of its three parts were non-textual… See the full description on the dataset page: https://huggingface.co/datasets/timodonnell/bioreason-pro-sft-reasoning-documents.tabulartext-generation100K<n<1M0 likes57 downloads2mo agoHugging Face30Piyushmryaa /california-housing-documentedtabulartabular-regression10K<n<100K0 likes52 downloads25d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.