CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abigailhaddad /foia-reading-room-documents Foia Reading Room Documents Documents from federal FOIA reading rooms and Inspector General report libraries: audits, inspections, investigative summaries and records released under the Freedom of Information Act. Every document here was published by a US federal agency and is a work of the United States government. Nothing has been altered: files are byte-identical to what the agency posted, and the checksum in metadata.parquet is of the original bytes. Why this… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/foia-reading-room-documents.text-retrieval0 likes5.8k downloads9d agoHugging Face02AiAF /JFK-Assassination-Records-2025-Documents-Release0 likes3.4k downloads1y agoHugging Face03mtybilly /apex-r1-real-world-documents Apex-R1 Real-World Benchmark Documents This dataset stores real-world document/data assets collected for Apex-R1 synthetic long-horizon agentic RL workspace generation. The files are intended as seed workspace materials, not as benchmark task labels. They can be injected into APEX-style filesystem/ or .apps_data/ environments to create more realistic and diverse professional-domain tasks. Contents benchmark_documents/ EnterpriseBench/ # CRM invoices… See the full description on the dataset page: https://huggingface.co/datasets/mtybilly/apex-r1-real-world-documents.documentdocument-question-answeringn<1K0 likes2.5k downloads2mo agoHugging Face04sunnysetia /locus-pretrain-documents-public Locus pretraining: documents public Setup only. No corpus release has been published and no canary has been run for this format. This repository will hold versioned indexed token artifacts. Tokens are little-endian int32 binary arrays; signed int64 offsets identify complete documents or unpadded pieces. Loss masks are bit-packed, with Parquet indexes and provenance sidecars. Training packs pieces using RSDB. Only a complete, verified manifest pinned to an immutable commit… See the full description on the dataset page: https://huggingface.co/datasets/sunnysetia/locus-pretrain-documents-public.0 likes1.8k downloads13d agoHugging Face05SherlockRamos /jurisdb-legal-documents JurisDB - Brazilian Legal Documents Dataset Dataset Description This dataset contains a comprehensive collection of Brazilian legal documents, including legislation (federal and state laws) and jurisprudence (court rulings and summaries from TSE, STJ, STF, and TNU). Dataset Structure . ├── legislacao_grifada_e_anotada_atualiz_em_01_01_2026/ │ ├── leis_estaduais/ │ ├── leis_federais/ │ └── ... └── sumulas_tse_stj_stf_e_tnu_atualiz_01_01_2026_2/ ├──… See the full description on the dataset page: https://huggingface.co/datasets/SherlockRamos/jurisdb-legal-documents.documenttext-classificationn<1K0 likes1.7k downloads8mo agoHugging Face06CGIAR /Embrapa-ai-documents-markdown1 likes1.7k downloads2mo agoHugging Face07Voxel51 /form_understanding_in_noisy_scanned_documents_plus Dataset Card for Form Understanding in Noisy Scanned Documents Plus This is a FiftyOne dataset with 1026 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("Voxel51/form_understanding_in_noisy_scanned_documents_plus") # Launch the App… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/form_understanding_in_noisy_scanned_documents_plus.imageobject-detection1K<n<10K1 likes1.6k downloads11mo agoHugging Face08abigailhaddad /sam-solicitation-documents Sam Solicitation Documents Attachments from federal solicitation notices on SAM.gov: statements of work, performance work statements, justifications, amendments, wage determinations and the rest of the paperwork that accompanies a federal contract opportunity. Every document here was published by a US federal agency and is a work of the United States government. Nothing has been altered: files are byte-identical to what the agency posted, and the checksum in metadata.parquet is… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/sam-solicitation-documents.text-retrieval0 likes1.3k downloads5d agoHugging Face09HumynLabs /Arabic_Documents_Dataset_PDF Arabic Documents Dataset (PDF) This dataset contains a collection of Arabic-language documents in PDF format. The corpus includes books, articles, reports, and educational materials written in Modern Standard Arabic and regional variants. It is curated to support AI research in document understanding, Arabic OCR, and text extraction from complex layouts. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Arabic_Documents_Dataset_PDF.documentn<1K0 likes1.2k downloads11mo agoHugging Face10HumynLabs /French_Documents_Dataset_PDF French Documents Dataset (PDF) This dataset contains a curated collection of French-language documents in PDF format. It includes educational materials, books, news articles, government publications, and public-domain literature written in French. The dataset supports AI research in OCR, document understanding, and multilingual text extraction. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/French_Documents_Dataset_PDF.documentn<1K0 likes1.2k downloads11mo agoHugging Face11singletongue /cc100-documents cc100-documents This dataset is a restructured version of the CC-100 (statmt/cc100) dataset. In the original dataset, each instance corresponds to a single paragraph (or a document boundary). In this version, the data has been reformed so that each instance corresponds to a single, complete document. This document-level structure makes it more convenient for processing using the map() and filter() methods in the Hugging Face Datasets library. Languages The following… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/cc100-documents.tabulartext-generation100M<n<1B1 likes1.1k downloads11mo agoHugging Face12th1nhng0 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.). Curated by: Thịnh Ngô Source: vbpl.vn… See the full description on the dataset page: https://huggingface.co/datasets/th1nhng0/vietnamese-legal-documents.texttext-classification1M<n<10M45 likes924 downloads2mo agoHugging Face13nyuuzyou /znanio-documents Dataset Card for Znanio.ru Educational Documents Dataset Summary This dataset contains 588,545 educational document files from the znanio.ru platform, a resource for teachers, educators, students, and parents providing diverse educational content. Znanio.ru has been a pioneer in educational technologies and distance learning in the Russian-speaking internet since 2009. The dataset includes a small portion of English language content, primarily for language learning… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/znanio-documents.texttext-classification100K<n<1M0 likes889 downloads2y agoHugging Face14sentence-transformers /example-documents Example Documents A small set of example documents across modalities (image, audio, video) for use in Sentence Transformers retrieval snippets and documentation. These are the kinds of files you pass to model.encode_document(...). They can safely be used as examples in your model cards if you don't want to host the example assets in your model repositories themselves. Contents File Modality doc1.jpg image (document page) doc2.jpg image (document page)… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/example-documents.audion<1K1 likes868 downloads1mo agoHugging Face15jinaai /wikimedia-commons-documents-ml_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_beir.image1K<n<10K0 likes804 downloads1y agoHugging Face16CGIAR /ifpri-ai-documents-markdown GAIA / GARDIAN-CIGI Agricultural Research Corpus This dataset contains 21,726 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline. Dataset Overview Property Value Total Documents 21,726 Total Size 623.27 MB Total Tokens 85,359,442 Total Pages 0 Languages 25 Unique Keywords 7,127 Resource Types 20 Date Generated 2026-07-31 02:55:20 Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/ifpri-ai-documents-markdown.summarization10K<n<100K1 likes798 downloads2mo agoHugging Face17DBWBD /Chinese_Debate_Documents Dataset Card for Chinese Debate Documents ASR-transcribed corpus of competitive Mandarin Chinese university debates, speaker-segmented and timestamped, with topic / round / team metadata parsed from the source filenames. Loading the Dataset from datasets import load_dataset ds = load_dataset("DBWBD/Chinese_Debate_Documents", split="train") print(ds[0]["topic"], "—", ds[0]["team_a"], "vs", ds[0]["team_b"]) for seg in ds[0]["segments"][:3]: print(f"… See the full description on the dataset page: https://huggingface.co/datasets/DBWBD/Chinese_Debate_Documents.tabulartext-classification1K<n<10K1 likes766 downloads4mo agoHugging Face18HumynLabs /Italian_Documents_Dataset_PDF Italian Documents Dataset (PDF) This dataset contains a curated collection of Italian-language documents in PDF format. It includes books, academic publications, reports, government documents, and news articles written in Italian. The dataset supports AI research in OCR, multilingual document understanding, and text recognition for Romance languages. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Italian_Documents_Dataset_PDF.documentn<1K0 likes724 downloads11mo agoHugging Face19HumynLabs /German_Documents_Dataset_PDFdocumentn<1K0 likes709 downloads11mo agoHugging Face20hotchpotch /cc100-ja-documents cc100-ja-documents HuggingFace で公開されている cc100 / cc100-ja は line 単位の分割のため、document 単位に結合したものです。 ライセンスはオリジナルのcc100 に準拠します。 text10M<n<100M4 likes670 downloads2y agoHugging Face21HumynLabs /Japanese_Documents_Dataset_PDF Japanese Documents Dataset (PDF) This dataset contains a curated collection of Japanese-language documents in PDF format. The corpus includes textbooks, research papers, news articles, public-domain books, and government publications written in Japanese. It is intended to support AI research in OCR, document understanding, and multilingual text recognition. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Japanese_Documents_Dataset_PDF.documentn<1K1 likes666 downloads11mo agoHugging Face22HumynLabs /Chinese_Documents_Dataset_PDF Chinese Documents Dataset (PDF) This dataset consists of a curated collection of Chinese-language documents in PDF format. It includes textbooks, research papers, articles, public-domain books, and official documents written in Simplified and Traditional Chinese. The dataset supports AI research in OCR, document understanding, and multilingual text extraction. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Chinese_Documents_Dataset_PDF.documentn<1K0 likes656 downloads11mo agoHugging Face23DeyangKong /documents_QA_100docs0 likes647 downloads2y agoHugging Face24abigailhaddad /govinfo-documents Govinfo Documents An index of what the Government Publishing Office publishes on govinfo.gov -- congressional hearings, committee reports and prints, congressional documents, GAO reports, agency publications and presidential documents. One row per document with title, agency, date, page count, checksum and the URL the PDF is served from. The files themselves are not mirrored here: GPO guarantees permanent public access to them, so copying them would duplicate a corpus that is… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/govinfo-documents.text-retrieval0 likes645 downloads12d agoHugging Face25CGIAR /gardian-cigi-ai-documents-markdown1 likes632 downloads2mo agoHugging Face26ZamAI-Pashto /zamai-pashto-documents ZamAI-Pashto Documents This repository contains the dataset and scripts for the ZamAI-Pashto document processing project, focusing exclusively on the Pashto language. Project Structure data/: Contains scanned documents, extracted text, translations, and summaries. annotations/: OCR bounding boxes, handwriting labels, and domain tags. scripts/: OCR processing, text cleaning, and translation alignment scripts. configs/: Dataset configurations and OCR settings.… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-documents.textvisual-document-retrievaln<1K0 likes612 downloads2mo agoHugging Face27hulk10 /legi-full-documentstext100K<n<1M0 likes611 downloads17h agoHugging Face28INSAIT-Institute /SaberMath-documents SABER-Math Documents Document corpus for the SABER-Math mathematical information reranking benchmark. Each of the 71,117 entries is a competition-style math problem together with its solution, serving as a retrievable document. The candidates field in SABER-Math Queries indexes into this corpus. Dataset structure Each example contains: Field Type… See the full description on the dataset page: https://huggingface.co/datasets/INSAIT-Institute/SaberMath-documents.text-retrieval10K<n<100K0 likes584 downloads1mo agoHugging Face29ImranzamanML /Clinical_Documents_on_Syndromes_Diseasetext10K<n<100K3 likes568 downloads2y agoHugging Face30HumynLabs /Russian_Documents_Dataset_PDF Russian Documents Dataset (PDF) This dataset contains a curated collection of Russian-language documents in PDF format. The corpus includes books, academic papers, government publications, articles, and educational materials written in Russian. It is designed to support AI research in OCR, document understanding, and multilingual text recognition. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Russian_Documents_Dataset_PDF.documentn<1K1 likes532 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.