CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01huggingface /documentation-images This dataset contains images used in the documentation of HuggingFace's libraries. HF Team: Please make sure you optimize the assets before uploading them. My favorite tool for this is https://tinypng.com/. imagen<1K204 likes2m downloads10h agoHugging Face02huggingface-course /documentation-imagesimagen<1K3 likes306k downloads1y agoHugging Face03AmazonScience /document-haystack Document Haystack Dataset This repository contains the dataset for the paper “Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark”. 📑 Abstract Paper The proliferation of multimodal Large Language Models has significantly advanced the ability to analyze and understand complex data inputs from different modalities. However, the processing of long documents remains under-explored, largely due to a lack of suitable benchmarks. To… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/document-haystack.textquestion-answering20 likes89k downloads1y agoHugging Face04EleutherAI /wikitext_document_level Wikitext Document Level This is a modified version of https://huggingface.co/datasets/wikitext that returns Wiki pages instead of Wiki text line-by-line. The original readme is contained below. Dataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/wikitext_document_level.text10K<n<100K18 likes76k downloads2y agoHugging Face05trl-lib /documentation-imagesimagen<1K0 likes47k downloads19d agoHugging Face06pruna-test /documentation-mediaimagen<1K0 likes35k downloads7d agoHugging Face07jinaai /documentation-imagesimagen<1K0 likes28k downloads1y agoHugging Face08tiiuae /documentation-imagesimagen<1K0 likes7.4k downloads1y agoHugging Face09optimum /documentation-imagesThis dataset contains images used in the documentation of HuggingFace's Optimum library. imagen<1K2 likes6.4k downloads2mo agoHugging Face10abigailhaddad /foia-reading-room-documents Foia Reading Room Documents Documents from federal FOIA reading rooms and Inspector General report libraries: audits, inspections, investigative summaries and records released under the Freedom of Information Act. Every document here was published by a US federal agency and is a work of the United States government. Nothing has been altered: files are byte-identical to what the agency posted, and the checksum in metadata.parquet is of the original bytes. Why this… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/foia-reading-room-documents.text-retrieval0 likes5.8k downloads9d agoHugging Face11AiAF /JFK-Assassination-Records-2025-Documents-Release0 likes3.4k downloads1y agoHugging Face12ybelkada /documentation-imagesimagen<1K0 likes3.1k downloads3y agoHugging Face13Chunte /documentation-imagesimagen<1K0 likes3.1k downloads4mo agoHugging Face14HuggingFaceM4 /DocumentVQAimage10K<n<100K46 likes3k downloads3y agoHugging Face15mannycooper /document-review-data Document Review Data Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package. Current Title Extraction Dataset Surface Canonical prefix: datasets/title_extraction/ Effective datasets: datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/ datasets/title_extraction/evaluation/real_device_280_v1/ datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/ The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.tabular100K<n<1M3 likes3k downloads11d agoHugging Face16mtybilly /apex-r1-real-world-documents Apex-R1 Real-World Benchmark Documents This dataset stores real-world document/data assets collected for Apex-R1 synthetic long-horizon agentic RL workspace generation. The files are intended as seed workspace materials, not as benchmark task labels. They can be injected into APEX-style filesystem/ or .apps_data/ environments to create more realistic and diverse professional-domain tasks. Contents benchmark_documents/ EnterpriseBench/ # CRM invoices… See the full description on the dataset page: https://huggingface.co/datasets/mtybilly/apex-r1-real-world-documents.documentdocument-question-answeringn<1K0 likes2.5k downloads2mo agoHugging Face17embedl /documentation-imagesimagen<1K0 likes2.4k downloads2mo agoHugging Face18sunnysetia /locus-pretrain-documents-public Locus pretraining: documents public Setup only. No corpus release has been published and no canary has been run for this format. This repository will hold versioned indexed token artifacts. Tokens are little-endian int32 binary arrays; signed int64 offsets identify complete documents or unpadded pieces. Loss masks are bit-packed, with Parquet indexes and provenance sidecars. Training packs pieces using RSDB. Only a complete, verified manifest pinned to an immutable commit… See the full description on the dataset page: https://huggingface.co/datasets/sunnysetia/locus-pretrain-documents-public.0 likes1.8k downloads13d agoHugging Face19RoboCOIN /AI2_Alphabot_2_stamp_document AI2_Alphabot_2_stamp_document Dataset Description This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Task Preview View Video Directly Overview Total Episodes: 987 Total Frames: 369702 FPS: 30 Dataset Size: 7.18 GB Robot Name: AI2_Alphabot_2 End-Effector Type: two_finger_end_effector Teleoperation Type: vr_controller Sensors: cam_front_chest_rgb, cam_front_head_rgb, cam_left_wrist_rgb… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AI2_Alphabot_2_stamp_document.robotics0 likes1.8k downloads3mo agoHugging Face20SherlockRamos /jurisdb-legal-documents JurisDB - Brazilian Legal Documents Dataset Dataset Description This dataset contains a comprehensive collection of Brazilian legal documents, including legislation (federal and state laws) and jurisprudence (court rulings and summaries from TSE, STJ, STF, and TNU). Dataset Structure . ├── legislacao_grifada_e_anotada_atualiz_em_01_01_2026/ │ ├── leis_estaduais/ │ ├── leis_federais/ │ └── ... └── sumulas_tse_stj_stf_e_tnu_atualiz_01_01_2026_2/ ├──… See the full description on the dataset page: https://huggingface.co/datasets/SherlockRamos/jurisdb-legal-documents.documenttext-classificationn<1K0 likes1.7k downloads8mo agoHugging Face21CGIAR /Embrapa-ai-documents-markdown1 likes1.7k downloads2mo agoHugging Face22Voxel51 /form_understanding_in_noisy_scanned_documents_plus Dataset Card for Form Understanding in Noisy Scanned Documents Plus This is a FiftyOne dataset with 1026 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("Voxel51/form_understanding_in_noisy_scanned_documents_plus") # Launch the App… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/form_understanding_in_noisy_scanned_documents_plus.imageobject-detection1K<n<10K1 likes1.6k downloads11mo agoHugging Face23NLPC-UOM /document_alignment_dataset-Sinhala-Tamil-English Dataset summary This is a gold-standard benchmark dataset for document alignment, between Sinhala-English-Tamil languages. Data had been crawled from the following news websites. News Source url Army https://www.army.lk/ Hiru http://www.hirunews.lk ITN https://www.newsfirst.lk Newsfirst https://www.itnnews.lk The aligned documents have been manually annotated. Dataset The folder structure for each news source is as follows. army |--Sinhala… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/document_alignment_dataset-Sinhala-Tamil-English.sentence-similarity2 likes1.5k downloads3y agoHugging Face24Voxel51 /document-haystack-10pages Dataset Card for document-haystack-10pages This is a FiftyOne dataset with 250 samples. It's the 10-page subset of the full dataset. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("Voxel51/document-haystack-10pages") # Launch the App… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/document-haystack-10pages.imageimage-classificationn<1K1 likes1.4k downloads11mo agoHugging Face25abigailhaddad /sam-solicitation-documents Sam Solicitation Documents Attachments from federal solicitation notices on SAM.gov: statements of work, performance work statements, justifications, amendments, wage determinations and the rest of the paperwork that accompanies a federal contract opportunity. Every document here was published by a US federal agency and is a work of the United States government. Nothing has been altered: files are byte-identical to what the agency posted, and the checksum in metadata.parquet is… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/sam-solicitation-documents.text-retrieval0 likes1.3k downloads5d agoHugging Face26Upabjojr /documenti-societari-italiani-rag-eval0 likes1.3k downloads2y agoHugging Face27eustlb /documentation-imagesimagen<1K0 likes1.2k downloads10mo agoHugging Face28HumynLabs /Arabic_Documents_Dataset_PDF Arabic Documents Dataset (PDF) This dataset contains a collection of Arabic-language documents in PDF format. The corpus includes books, articles, reports, and educational materials written in Modern Standard Arabic and regional variants. It is curated to support AI research in document understanding, Arabic OCR, and text extraction from complex layouts. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Arabic_Documents_Dataset_PDF.documentn<1K0 likes1.2k downloads11mo agoHugging Face29RoboCOIN /Agilex_Cobot_Magic_zip_up_the_document_baggated Agilex_Cobot_Magic_zip_up_the_document_bag 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: Agilex_Cobot_Magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: office 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pull place 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_zip_up_the_document_bag.tabularrobotics100K<n<1M0 likes1.2k downloads3mo agoHugging Face30HumynLabs /French_Documents_Dataset_PDF French Documents Dataset (PDF) This dataset contains a curated collection of French-language documents in PDF format. It includes educational materials, books, news articles, government publications, and public-domain literature written in French. The dataset supports AI research in OCR, document understanding, and multilingual text extraction. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/French_Documents_Dataset_PDF.documentn<1K0 likes1.2k downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.