CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ZamAI-Pashto /zamai-pashto-documents ZamAI-Pashto Documents This repository contains the dataset and scripts for the ZamAI-Pashto document processing project, focusing exclusively on the Pashto language. Project Structure data/: Contains scanned documents, extracted text, translations, and summaries. annotations/: OCR bounding boxes, handwriting labels, and domain tags. scripts/: OCR processing, text cleaning, and translation alignment scripts. configs/: Dataset configurations and OCR settings.… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-documents.textvisual-document-retrievaln<1K0 likes612 downloads2mo agoHugging Face02jiwoochris /easylaw_kr_documentstext1K<n<10K2 likes58 downloads3y agoHugging Face03flaviawallen /MNLP_M3_rag_documentstext1K<n<10K0 likes52 downloads1y agoHugging Face04QIRIM /crh-parallel-corpora-document-level-noisytabulartranslation10K<n<100K1 likes51 downloads2y agoHugging Face05ClarusC64 /legal-privilege-log-document-basis-waiver-risk-v0.1What this dataset does You receive doc description date author recipients privilege basis redaction choice context waiver flags You decide coherent or incoherent Daily use privilege log QC waiver risk detection disclosure challenge prep tabulartext-classificationn<1K0 likes49 downloads7mo agoHugging Face06ClarusC64 /maritime-bill-of-lading-document-set-coherence-risk-v0.1What this repo is for Triage trade doc packs before they trigger holds. You use it to flag HS code inconsistencies across documents missing certificates shipper or consignee mismatch clearance status lag not supported by doc quality Why it matters Most port delay disputes begin in paperwork. texttext-classificationn<1K1 likes49 downloads7mo agoHugging Face07Nucleo360 /documentos-laborales-obligatorios-espana Documentos laborales de entrega obligatoria en España 14 documentos que la normativa laboral española obliga a elaborar o entregar, con a quién se entregan, en qué plazo, cuánto hay que conservarlos y el artículo concreto que lo exige. Publicado por Nucleo360, software de recursos humanos para pymes españolas. Por qué este conjunto de datos En materia laboral, el problema rara vez es no tener un documento: es no poder demostrar que se entregó. La mayoría de las… See the full description on the dataset page: https://huggingface.co/datasets/Nucleo360/documentos-laborales-obligatorios-espana.texttable-question-answeringn<1K0 likes37 downloads20d agoHugging Face08itsjhuang /watsonx-docs-document-type Watsonx Docs Document Type Classification This dataset is a balanced binary document-level classification subset derived from ibm-research/watsonxDocsQA. Task Classify IBM Watsonx documentation pages by their dominant user-facing purpose: conceptual: documents primarily used to understand or look up information. how-to: documents primarily used to complete a procedure or fix a problem. Splits Split conceptual how-to Total train 140 140 280… See the full description on the dataset page: https://huggingface.co/datasets/itsjhuang/watsonx-docs-document-type.texttext-classificationn<1K0 likes33 downloads5mo agoHugging Face09passport-visa-photo-studio /document-photo-requirements Verified Document Photo Requirements Dataset A structured reference collection of official-source passport, visa, and national ID photo requirements maintained by Passport Visa Photo Studio. It is not a photo corpus, training dataset, or model artifact. Dataset summary Version: 1.0.0 Release date: 2026-08-14 Latest source review represented: 2026-08-11 Records: 18 (12 passport, 5 visa, 1 national ID) Coverage: 14 countries or regions Formats: CSV and JSON… See the full description on the dataset page: https://huggingface.co/datasets/passport-visa-photo-studio/document-photo-requirements.tabularn<1K0 likes32 downloads1mo agoHugging Face10m-ric /transformers_documentation_entextn<1K0 likes26 downloads3y agoHugging Face11tasal9 /zamai-pashto-documents Pashto Languages: psLicense: cc-by-4.0Task categories: visual-document-retrievalSize categories: n<1K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for visual-document-retrieval tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/zamai-pashto-documents") print(dataset) Citation @misc{zamai_pashto_data, title = {{Pashto}}, author = {ZamAI / Yaqoob… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/zamai-pashto-documents.textvisual-document-retrievaln<1K0 likes26 downloads2mo agoHugging Face12Keboola /Developer-Documentation-QAtextquestion-answering1K<n<10K0 likes23 downloads2y agoHugging Face13ClarusC64 /legal-document-version-redline-final-coherence-risk-v0.1What this dataset does You receive version history redline summary final id sent or filed id approval record mismatch flags You decide coherent or incoherent Daily use wrong attachment prevention filing version QC approval gap detection tabulartext-classificationn<1K0 likes22 downloads7mo agoHugging Face14hoanglvuit /Legal-Documenttext100K<n<1M0 likes21 downloads9mo agoHugging Face15YuITC /vietnam-legal-documentstext100K<n<1M1 likes21 downloads5mo agoHugging Face1652100303-TranPhuocSang /vi-legal-documents-20k-pretraintext10K<n<100K0 likes20 downloads2y agoHugging Face17mtyrrell /NDC_documents_mastertabularn<1K0 likes19 downloads3y agoHugging Face18ClarusC64 /legal-chronology-event-document-issue-coherence-risk-v0.1What this dataset does You receive timeline summary document map issue links date checks gap flags conflict flags You decide coherent or incoherent Daily use chronology QC date conflict detection missing evidence detection gap finding tabulartext-classificationn<1K0 likes19 downloads7mo agoHugging Face19pranay27sy /maritime_documents_tags_classificationData Understanding and Preparation Data Collection BIMCO Contracts and Clauses URL: BIMCO Contracts and Clauses Description: This website provides a wide range of standardized contracts and clauses commonly used in the maritime industry. The documents were downloaded and used as part of our dataset, offering detailed insights into industry-specific terminology and structured data. Data Description Documents Collected: A total of 217 documents were collected… See the full description on the dataset page: https://huggingface.co/datasets/pranay27sy/maritime_documents_tags_classification.texttext-classification1K<n<10K1 likes17 downloads2y agoHugging Face20yjching /sas-documentation-v1textn<1K0 likes16 downloads3y agoHugging Face21sriramahesh2000 /DocumentCreationtextn<1K1 likes16 downloads3y agoHugging Face22abhinavdread /msme-dispute-document-corpus MSME Dispute Document Corpus (Synthetic OCR) Dataset Description This dataset contains 8,000+ synthetic document samples designed to train AI models for the Indian MSME (Micro, Small, and Medium Enterprises) dispute resolution sector. It is specifically engineered to handle Real-World OCR Noise and Adversarial Edge Cases (e.g., distinguishing a "Proforma Invoice" from a valid "Tax Invoice"). The data mimics the messy, unstructured text often found in scanned PDFs, photos… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/msme-dispute-document-corpus.tabulartext-classification1K<n<10K0 likes16 downloads7mo agoHugging Face23ClarusC64 /clinical-care-plan-documentation-action-coherence-risk-v0.1What this repo is for Detect when a plan exists in notes but orders or actions do not follow Examples you can use “CT today” documented but not ordered “stop heparin” documented but continues “review cultures” documented but never done You use it to flag care drift risk before harm texttext-classificationn<1K0 likes16 downloads7mo agoHugging Face24jason1966 /adwaittagalpallewar_medical-document-ocr-text-dataset Medical Document OCR Text Dataset Synthetic OCR-extracted text from medical documents for NLP classification Dataset Info Source: Kaggle Original Size: 22.39 MB Kaggle Downloads: 33 Files: 1 Files medical_documents_dataset.csv Mirrored from Kaggle text10K<n<100K0 likes16 downloads6mo agoHugging Face25Hadisawara /indonesian-tax-document-classification Indonesian Tax Document Classification Dataset Dataset Description Dataset ini berisi koleksi sintetis dokumen pajak Indonesia yang digunakan untuk klasifikasi jenis dokumen pajak. Dataset dirancang untuk mendukung penelitian NLP berbahasa Indonesia di bidang administrasi pajak dan pemerintahan daerah. Dataset ini dibuat berdasarkan pengalaman dan pengetahuan dari sistem administrasi pajak daerah (Bapenda), dengan struktur yang mencerminkan dokumen-dokumen nyata… See the full description on the dataset page: https://huggingface.co/datasets/Hadisawara/indonesian-tax-document-classification.tabulartext-classification10K<n<100K0 likes16 downloads3mo agoHugging Face26Prakhar1000 /API_Documentation_dataset_alpaancotextn<1K1 likes15 downloads3y agoHugging Face27ushakov15 /MNLP_M2_rag_documentstext10K<n<100K0 likes15 downloads1y agoHugging Face28abhinavdread /msme-document-presence-dataset MSME Document Presence Detection Dataset Overview This dataset is designed for training binary classification models to detect the presence of mandatory documents in MSME arbitration cases using OCR-extracted text. The dataset supports automated document completeness validation systems. Each sample represents a structured arbitration case with document-specific OCR text fields and binary presence labels. Documents Covered The dataset includes detection labels… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/msme-document-presence-dataset.tabulartext-classification10K<n<100K0 likes15 downloads7mo agoHugging Face29anasse15 /MNLP_M3_rag_documenttext10K<n<100K0 likes14 downloads1y agoHugging Face30hotamago /SoICT-Hackathon-2024-Legal-Document-Retrievaltext100K<n<1M0 likes13 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.