datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zamai-pashto-documents
ZamAI-Pashto Documents
This repository contains the dataset and scripts for the ZamAI-Pashto document processing project, focusing exclusively on the Pashto language.
Project Structure
data/: Contains scanned documents, extracted text, translations, and summaries.
annotations/: OCR bounding boxes, handwriting labels, and domain tags.
scripts/: OCR processing, text cleaning, and translation alignment scripts.
configs/: Dataset configurations and OCR settings.… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-documents.easylaw_kr_documentsMNLP_M3_rag_documentscrh-parallel-corpora-document-level-noisylegal-privilege-log-document-basis-waiver-risk-v0.1What this dataset does
You receive
doc description
date
author
recipients
privilege basis
redaction choice
context
waiver flags
You decide
coherent
or
incoherent
Daily use
privilege log QC
waiver risk detection
disclosure challenge prep
maritime-bill-of-lading-document-set-coherence-risk-v0.1What this repo is for
Triage trade doc packs before they trigger holds.
You use it to flag
HS code inconsistencies across documents
missing certificates
shipper or consignee mismatch
clearance status lag not supported by doc quality
Why it matters
Most port delay disputes begin in paperwork.
documentos-laborales-obligatorios-espana
Documentos laborales de entrega obligatoria en España
14 documentos que la normativa laboral española obliga a elaborar o entregar, con a quién se entregan, en qué plazo, cuánto hay que conservarlos y el artículo concreto que lo exige.
Publicado por Nucleo360, software de recursos humanos para pymes españolas.
Por qué este conjunto de datos
En materia laboral, el problema rara vez es no tener un documento: es no poder demostrar que se entregó. La mayoría de las… See the full description on the dataset page: https://huggingface.co/datasets/Nucleo360/documentos-laborales-obligatorios-espana.watsonx-docs-document-type
Watsonx Docs Document Type Classification
This dataset is a balanced binary document-level classification subset derived
from ibm-research/watsonxDocsQA.
Task
Classify IBM Watsonx documentation pages by their dominant user-facing purpose:
conceptual: documents primarily used to understand or look up information.
how-to: documents primarily used to complete a procedure or fix a problem.
Splits
Split
conceptual
how-to
Total
train
140
140
280… See the full description on the dataset page: https://huggingface.co/datasets/itsjhuang/watsonx-docs-document-type.document-photo-requirements
Verified Document Photo Requirements Dataset
A structured reference collection of official-source passport, visa, and national ID photo requirements maintained by Passport Visa Photo Studio. It is not a photo corpus, training dataset, or model artifact.
Dataset summary
Version: 1.0.0
Release date: 2026-08-14
Latest source review represented: 2026-08-11
Records: 18 (12 passport, 5 visa, 1 national ID)
Coverage: 14 countries or regions
Formats: CSV and JSON… See the full description on the dataset page: https://huggingface.co/datasets/passport-visa-photo-studio/document-photo-requirements.transformers_documentation_enzamai-pashto-documents
Pashto
Languages: psLicense: cc-by-4.0Task categories: visual-document-retrievalSize categories: n<1K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for visual-document-retrieval tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/zamai-pashto-documents")
print(dataset)
Citation
@misc{zamai_pashto_data,
title = {{Pashto}},
author = {ZamAI / Yaqoob… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/zamai-pashto-documents.Developer-Documentation-QAlegal-document-version-redline-final-coherence-risk-v0.1What this dataset does
You receive
version history
redline summary
final id
sent or filed id
approval record
mismatch flags
You decide
coherent
or
incoherent
Daily use
wrong attachment prevention
filing version QC
approval gap detection
Legal-Documentvietnam-legal-documentsvi-legal-documents-20k-pretrainNDC_documents_masterlegal-chronology-event-document-issue-coherence-risk-v0.1What this dataset does
You receive
timeline summary
document map
issue links
date checks
gap flags
conflict flags
You decide
coherent
or
incoherent
Daily use
chronology QC
date conflict detection
missing evidence detection
gap finding
maritime_documents_tags_classificationData Understanding and Preparation
Data Collection
BIMCO Contracts and Clauses
URL: BIMCO Contracts and Clauses
Description: This website provides a wide range of standardized contracts and clauses commonly used in the maritime industry. The documents were downloaded and used as part of our dataset, offering detailed insights into industry-specific terminology and structured data.
Data Description
Documents Collected: A total of 217 documents were collected… See the full description on the dataset page: https://huggingface.co/datasets/pranay27sy/maritime_documents_tags_classification.sas-documentation-v1DocumentCreationmsme-dispute-document-corpus
MSME Dispute Document Corpus (Synthetic OCR)
Dataset Description
This dataset contains 8,000+ synthetic document samples designed to train AI models for the Indian MSME (Micro, Small, and Medium Enterprises) dispute resolution sector.
It is specifically engineered to handle Real-World OCR Noise and Adversarial Edge Cases (e.g., distinguishing a "Proforma Invoice" from a valid "Tax Invoice"). The data mimics the messy, unstructured text often found in scanned PDFs, photos… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/msme-dispute-document-corpus.clinical-care-plan-documentation-action-coherence-risk-v0.1What this repo is for
Detect when
a plan exists in notes
but orders
or actions
do not follow
Examples you can use
“CT today” documented but not ordered
“stop heparin” documented but continues
“review cultures” documented but never done
You use it to flag
care drift risk
before harm
adwaittagalpallewar_medical-document-ocr-text-dataset
Medical Document OCR Text Dataset
Synthetic OCR-extracted text from medical documents for NLP classification
Dataset Info
Source: Kaggle
Original Size: 22.39 MB
Kaggle Downloads: 33
Files: 1
Files
medical_documents_dataset.csv
Mirrored from Kaggle
indonesian-tax-document-classification
Indonesian Tax Document Classification Dataset
Dataset Description
Dataset ini berisi koleksi sintetis dokumen pajak Indonesia yang digunakan untuk klasifikasi jenis dokumen pajak. Dataset dirancang untuk mendukung penelitian NLP berbahasa Indonesia di bidang administrasi pajak dan pemerintahan daerah.
Dataset ini dibuat berdasarkan pengalaman dan pengetahuan dari sistem administrasi pajak daerah (Bapenda), dengan struktur yang mencerminkan dokumen-dokumen nyata… See the full description on the dataset page: https://huggingface.co/datasets/Hadisawara/indonesian-tax-document-classification.API_Documentation_dataset_alpaancoMNLP_M2_rag_documentsmsme-document-presence-dataset
MSME Document Presence Detection Dataset
Overview
This dataset is designed for training binary classification models to detect the presence of mandatory documents in MSME arbitration cases using OCR-extracted text.
The dataset supports automated document completeness validation systems.
Each sample represents a structured arbitration case with document-specific OCR text fields and binary presence labels.
Documents Covered
The dataset includes detection labels… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/msme-document-presence-dataset.MNLP_M3_rag_documentSoICT-Hackathon-2024-Legal-Document-Retrieval
