information-extraction
rst-information-extraction-11bbpmn-information-extraction-v2donut-base-Medical_Handwritten_Prescriptions_Information_Extractiondonut-base-Medical_Handwritten_Prescriptions_Information_Extraction_1bpmn-information-extractionLayoutlm_Form_information_extractiondonut-base-Medical_Handwritten_Prescriptions_Information_Extraction_updateddonut-base-Medical_Handwritten_Prescriptions_Information_Extraction
poster-schedule-information-extraction
Multimodal Visual-Text Dataset for Poster Schedule Information Extraction
A ready-to-train Indonesian document AI dataset combining pixels, OCR tokens, spatial layout, and BIO entity labels.
Overview
What it is
127 Indonesian seminar and religious-study event posters with multimodal token-level annotations
Primary task
Schedule information extraction as token classification
Modalities
Image + text + 2D spatial layout
Coordinates… See the full description on the dataset page: https://huggingface.co/datasets/Ibnuck/poster-schedule-information-extraction.key_information_extractionreceipt_VLM_information_extractionhtml_to_json_information_extraction_dataset
HTML to JSON Information Extraction Dataset
Description
The html_to_json_information_extraction dataset is a collection of over 7300 HTML snippets and their extracted information in JSON.
These HTML have been sourced (scraped) from about 25 companies' career pages.
The dataset contains three splits - train, test, unseen_test.
This dataset has been built to fine tune SLMs & LLMs for the information extraction task.
train split
This split contains over 5700 pair… See the full description on the dataset page: https://huggingface.co/datasets/Jiraya/html_to_json_information_extraction_dataset.key-information-extraction
Neural Metrics · Straight at the task: pull the right fields out.
A key-information-extraction set aimed squarely at the core job - given a document, return the specific values that matter.
We use it for: benchmarking field-level precision and recall - measuring how often we hallucinate a plausible-but-absent value.
Attribution
This is an unmodified fork of nanonets/key_information_extraction, created by the Qwen team.
All weights, files and behaviour are… See the full description on the dataset page: https://huggingface.co/datasets/NeuralMetrics/key-information-extraction.kleister_nda_information_extraction
Kleister NDA — Information Extraction (orgrctera/kleister_nda_information_extraction)
Overview
This release packages the Kleister NDA split of the Kleister benchmark as rows suitable for information extraction (IE) evaluation. Each example points at a Non-Disclosure Agreement (NDA) document and specifies which attribute keys should be filled; the target is a JSON object of normalized string values for those keys (with null when a value is absent or not applicable).… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/kleister_nda_information_extraction.
