CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mannycooper /document-review-data Document Review Data Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package. Current Title Extraction Dataset Surface Canonical prefix: datasets/title_extraction/ Effective datasets: datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/ datasets/title_extraction/evaluation/real_device_280_v1/ datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/ The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.tabular100K<n<1M3 likes3k downloads11d agoHugging Face02VLM2Vec /MMLongBench-docimage10K<n<100K0 likes2.5k downloads1y agoHugging Face03timodonnell /protein-docs Protein Documents (Parquet) Structured text documents encoding protein residue sequences and 3D contact maps from AlphaFold Database v4 predicted structures, stored as Parquet files. Each row is one protein document with metadata. Source structures: timodonnell/afdb-24M and timodonnell/afdb-1.6M Document Schemes Each subdirectory contains documents generated with a different scheme. All schemes share leakage-resistant train/val/test splits based on structural… See the full description on the dataset page: https://huggingface.co/datasets/timodonnell/protein-docs.tabulartext-generation10M<n<100M0 likes2.1k downloads4mo agoHugging Face04vidore /docvqa_test_subsampled_beirBEIR version of vidore/docvqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes1.7k downloads1y agoHugging Face05doctor-ghezelbaash /dr-saeid-ghezelbaash-entity-data Dr. Saeed Ghezelbash Public Knowledge Graph A public, physician-authored knowledge graph and multilingual retrieval dataset by Dr. Saeed Ghezelbash, a physician in Kermanshah, Iran. It connects physician identity, aesthetic medicine services, published question-answer content and cited evidence for entity resolution and evidence-grounded AI retrieval. The canonical source is the official website and Dataset graph. This Hugging Face repository is its AI distribution. The… See the full description on the dataset page: https://huggingface.co/datasets/doctor-ghezelbaash/dr-saeid-ghezelbaash-entity-data.textquestion-answering1K<n<10K1 likes1.5k downloads23h agoHugging Face06ytzi /the-stack-dedup-python-filtered-docstringsThis is a dataset originated from bigcode/the-stack-dedup with some filters applied. The filters filtered in this dataset are: remove_function_no_docstring remove_class_no_docstring remove_delete_markers tabular10M<n<100M0 likes1.4k downloads2y agoHugging Face07Ayushnangia /docmath-eval-failures-200 DocMath-Eval Failures 200: Agent Benchmark & Leaderboard A curated benchmark of 200 challenging financial math questions that leading AI models failed to answer correctly, with comprehensive evaluation results from multiple AI agents. Leaderboard Evaluated on 2026-02-21 using LLM-as-Judge (Qwen QwQ-32B) for soft scoring. Rank Agent Model Exact Match Judge: Exact Judge: Approx Judge: Total Wrong Avg Duration Avg Tool Calls 1 TRAE Agent Opus 4.5 98/200 (49.0%) 96… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/docmath-eval-failures-200.tabularquestion-answering1K<n<10K0 likes1.2k downloads7mo agoHugging Face08RoboCOIN /Agilex_Cobot_Magic_zip_up_the_document_baggated Agilex_Cobot_Magic_zip_up_the_document_bag 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: Agilex_Cobot_Magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: office 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pull place 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_zip_up_the_document_bag.tabularrobotics100K<n<1M0 likes1.2k downloads3mo agoHugging Face09singletongue /cc100-documents cc100-documents This dataset is a restructured version of the CC-100 (statmt/cc100) dataset. In the original dataset, each instance corresponds to a single paragraph (or a document boundary). In this version, the data has been reformed so that each instance corresponds to a single, complete document. This document-level structure makes it more convenient for processing using the map() and filter() methods in the Hugging Face Datasets library. Languages The following… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/cc100-documents.tabulartext-generation100M<n<1B1 likes1.1k downloads11mo agoHugging Face10mteb /docvqa_test_subsampled_beirBEIR version of vidore/docvqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes1k downloads7mo agoHugging Face11eczech /marinfold-exp11-protein-docs-seq marinfold-exp11-pdocs-seq Sequence-only derivative of eczech/marinfold-exp11-protein-docs. For every row, the document field has been reduced to just the amino-acid sequence portion: the <begin_sequence> tag followed by the per-residue three-letter tokens (e.g. <begin_sequence> <MET> <LYS> <ASN> ...). The <contacts-and-distances-v1> document-type prefix and everything from <begin_statements> onward (contacts and distances) are removed. The token format is preserved verbatim so… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs-seq.tabulartext-generation1M<n<10M0 likes989 downloads4mo agoHugging Face12RoboCOIN /RMC-AIDA-L_organise_the_document_baggated RMC-AIDA-L_organise_the_document_bag 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: realman_rmc_aidal | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: office 🤖 Atomic Actions This dataset includes the following atomic actions: grasp place pick pull 📊 Dataset Statistics Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/RMC-AIDA-L_organise_the_document_bag.tabularrobotics100K<n<1M0 likes939 downloads9mo agoHugging Face13microsoft /XL-DocBench XL-DocBench Evidence-grounded reasoning across hundreds or thousands of pages. Fully verified by 194 human experts. Hongchen Wei1,†,‡, Yuanzhe Wang2,†,‡, Bei Liu2,*, Yifan Yang2, Qi Dai2, Ruichun Ma2, Kai Qiu2, Yunsheng Li2, Dongdong Chen2, Chong Luo2, Zhenzhong Chen1, Baining Guo2 1Wuhan University &nbsp; 2Microsoft &nbsp; †Equal contribution &nbsp; ‡Work done during an internship at MSRA &nbsp; *Project leader Project Page · Paper · Live Leaderboard… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/XL-DocBench.tabularquestion-answering1K<n<10K7 likes932 downloads20d agoHugging Face14eczech /marinfold-exp11-protein-docs marinfold-exp11-pdocs Quality-bucketed re-publication of the contacts-and-distances-v1-5x config from timodonnell/protein-docs, partitioned by the source round column: Config Source rounds Approx rows high round 0 ~1.68M medium round 1 ~1.42M low round 2–4 ~2.29M Train/val/test split assignment is inherited from the source dataset (leakage-resistant structural-cluster hashing). All columns from the source are preserved; rows are simply partitioned by round. See… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs.tabulartext-generation1M<n<10M0 likes874 downloads4mo agoHugging Face15r2e-edits /r2e-dockers-rllm-v1tabular10K<n<100K0 likes857 downloads1y agoHugging Face16jinaai /wikimedia-commons-documents-ml_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_beir.image1K<n<10K0 likes804 downloads1y agoHugging Face17DBWBD /Chinese_Debate_Documents Dataset Card for Chinese Debate Documents ASR-transcribed corpus of competitive Mandarin Chinese university debates, speaker-segmented and timestamped, with topic / round / team metadata parsed from the source filenames. Loading the Dataset from datasets import load_dataset ds = load_dataset("DBWBD/Chinese_Debate_Documents", split="train") print(ds[0]["topic"], "—", ds[0]["team_a"], "vs", ds[0]["team_b"]) for seg in ds[0]["segments"][:3]: print(f"… See the full description on the dataset page: https://huggingface.co/datasets/DBWBD/Chinese_Debate_Documents.tabulartext-classification1K<n<10K1 likes766 downloads4mo agoHugging Face18r2e-edits /dockersv1tabular1K<n<10K0 likes573 downloads2y agoHugging Face19ytzi /the-stack-dedup-python-filtered-docstrings-gpt2tabular10M<n<100M0 likes566 downloads2y agoHugging Face20semeru /text-code-galeras-code-generation-from-docstring-3k-dedupedtabular1K<n<10K0 likes565 downloads3y agoHugging Face21oklenAI /UDM_cleaned_docs UDM cleaned docs 6,029,052 web pages reduced to just their mathematical content, extracted verbatim by oklenAI/udm_doc_extract_qwen3.5_2B — a 2B model distilled from GPT-5.6. Every row is model output, not human-curated text. The extract field is what the model returned for that page; the source page text is not included. Read Two repetition flags below before filtering — the obvious flag is not the one you want. How it was built step pages… See the full description on the dataset page: https://huggingface.co/datasets/oklenAI/UDM_cleaned_docs.tabulartext-generation1M<n<10M0 likes562 downloads22d agoHugging Face22hulk10 /data_gouv_datasets_catalog-full-documents 🇫🇷 Catalogue des jeux de données de data.gouv.fr – Version structurée Ce dataset constitue une version structurée et exhaustive du catalogue des jeux de données publiés sur data.gouv.fr, la plateforme nationale française de l’open data. Il recense l’ensemble des jeux de données référencés sur la plateforme et fournit leurs métadonnées complètes : titre et description, organisation productrice, licence, couverture spatiale et temporelle, fréquence de mise à jour, formats… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/data_gouv_datasets_catalog-full-documents.tabular100K<n<1M0 likes526 downloads16h agoHugging Face23r2e-edits /r2e-dockers-v3tabular1K<n<10K0 likes524 downloads2y agoHugging Face24r2e-edits /r2e-dockers-v2tabular1K<n<10K0 likes519 downloads2y agoHugging Face25hulk10 /conseil-detat-full-documents Décisions du Conseil d'État (France) Description Ce dataset contient un corpus de décisions rendues par le Conseil d'État français, la plus haute juridiction de l'ordre administratif. Les décisions proviennent de la plateforme officielle Open Data de la Justice Administrative et sont diffusées au format XML anonymisé. Le corpus rassemble les textes intégraux des décisions ainsi que plusieurs métadonnées permettant leur identification et leur traçabilité. Source… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-detat-full-documents.tabularquestion-answering1K<n<10K1 likes513 downloads16h agoHugging Face26doctolib-lab /finemed-fr FineMed-fr 🤗 Blog | 📄 Paper | 💻 Code | 🌐 FineMed | 🩺 DoctoBERT 📚 Introduction FineMed-fr is a large, openly available corpus of French medical text for language-model pretraining: 21.1M documents and 19.2B words of real-world medical writing, annotated along several quality axes. The corpus is drawn from three heterogeneous open-web sources (FineWeb-2, FinePDFs, and FineWiki), which together provide the scale, source diversity, and stylistic range… See the full description on the dataset page: https://huggingface.co/datasets/doctolib-lab/finemed-fr.imagefill-mask10M<n<100M7 likes487 downloads3mo agoHugging Face27superdoc-dev /docx-corpus docx-corpus The largest classified corpus of Word documents. 736K+ .docx files from the public web, classified into 10 document types and 9 topics across 76 languages. Dataset Description This dataset contains metadata for publicly available .docx files collected from the web. Each document has been classified by document type and topic using a two-stage pipeline: LLM labeling (Claude) of a stratified sample, followed by fine-tuned XLM-RoBERTa classifiers applied at… See the full description on the dataset page: https://huggingface.co/datasets/superdoc-dev/docx-corpus.tabulartext-classification100K<n<1M5 likes372 downloads7mo agoHugging Face28Hieuman /gmane_dataset-doctabular1M<n<10M0 likes346 downloads1y agoHugging Face29zalizedata /us-court-opinions-dockets-judges-dataset US Court Opinions Metadata, Dockets & Judges (CourtListener) 10M opinion clusters, 70M dockets and 16K judges from official CourtListener / Free Law Project bulk data as metadata + derived-signals tables — citation graph, company litigation profiles; no opinion full text. Part of the DataForge Open Data program — full production packages, free for academic and personal use. Canonical dataset page: https://data.zalize.com/datasets/us-court-opinions-dockets-judges-dataset… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/us-court-opinions-dockets-judges-dataset.tabulartext-classification10M<n<100M0 likes341 downloads1mo agoHugging Face30Hizy /doc2 doc2lora_full_p0p1_text_v1 Text-format parquet dataset for document-to-LoRA (doc2lora) continual-learning pretraining. Plain-text fields (context/prompt/response/suffix) plus Qwen3.5-4B token id lists (*_tokens_qwen35) for arbitrary-model re-tokenization. Train split: train/ — 3287 parquet shards, 32,703,729 rows, ~7.7 GB. reconstruct_conversation: prompt + response + suffix ; supervised_text: response. Auto-generated card. Uploaded via hf-mirror.com. tabular1M<n<10M0 likes322 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.