datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
funsd-layoutlmv3SORIE_layoutlmv2
Dataset Card for "SORIE_layoutlmv2"
More Information needed
DocVQA_layoutLM
Dataset Card for "DocVQA_layoutLM"
More Information needed
DocVQA_for_LayoutLM
Dataset Card for "DocVQA_layoutLM_large"
More Information needed
layoutlmv3-invoice-dataset
LayoutLMv3 Invoice Dataset
This dataset is processed and ready for training LayoutLMv3 models for invoice information extraction.
Dataset Description
This dataset contains invoice documents with OCR-extracted text, bounding boxes, and entity labels for training document understanding models.
Dataset Structure
train: Training split
validation: Validation split (if available)
test: Test split (if available)
Features
input_ids: Tokenized text input… See the full description on the dataset page: https://huggingface.co/datasets/Kwash67/layoutlmv3-invoice-dataset.CIVQA-TesseractOCR-LayoutLM
CIVQA TesseractOCR LayoutLM Dataset
The Czech Invoice Visual Question Answering dataset was created with Tesseract OCR and encoded for the LayoutLM.
The pre-encoded dataset can be found on this link: https://huggingface.co/datasets/fimu-docproc-research/CIVQA-TesseractOCR
All invoices used in this dataset were obtained from public sources. Over these invoices, we were focusing on 15 different entities, which are crucial for processing the invoices.
Invoice number
Variable… See the full description on the dataset page: https://huggingface.co/datasets/SpringRollMonster/CIVQA-TesseractOCR-LayoutLM.layoutlm_dataset2xfund-multilingual-normalized-layoutlmv3layoutlmv3_cord
Dataset Card for "layoutlmv3_cord"
Original Dataset is "naver-clova-ix/cord-v2"
This dataset is modified for learning.
More Information needed
resume_parsing_layoutlmsroie-for-layoutlmarabic-books-layoutlm-formattedSROIE_layoutlmv2_sequence
Dataset Card for "SROIE_layoutlmv2_sequence"
More Information needed
layoutlmv3-cordv2-binary-mappedlayoutlmv3-cordv2layoutlmv3-cordv2-binaryinvoices_layoutlm_formatLayoutLMv3
Dataset Card for "LayoutLMv3"
More Information needed
funsd-layoutlmv3
Neural Metrics · Noisy scanned forms with key-value ground truth.
FUNSD is a small, deliberately grubby set of scanned forms annotated with entities and the links between them. It is the classic sanity check for form understanding.
We use it for: key-value pair extraction - entity linking on forms - checking that a model degrades gracefully on genuinely bad scans.
Attribution
This is an unmodified fork of nielsr/funsd-layoutlmv3, created by the Qwen team.… See the full description on the dataset page: https://huggingface.co/datasets/NeuralMetrics/funsd-layoutlmv3.layoutlmv3-finetuning-datafunsd-layoutlmv3LayoutLMv3-first
Dataset Card for "LayoutLMv3-first"
More Information needed
layoutlmv3_100_labelledlayoutlmv3layoutlm_sqad
Dataset Card for "layoutlm_sqad"
More Information needed
publaynet-layoutlmv3LayoutLM_data
Dataset Card for "LayoutLM_data"
More Information needed
layoutlmv3-document-qa-v2alayoutlmv3_employee_info_v1
Dataset Card for "layoutlmv3_employee_info_v1"
More Information needed
layoutlmv3-cordv2
