datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
poster-schedule-information-extraction
Multimodal Visual-Text Dataset for Poster Schedule Information Extraction
A ready-to-train Indonesian document AI dataset combining pixels, OCR tokens, spatial layout, and BIO entity labels.
Overview
What it is
127 Indonesian seminar and religious-study event posters with multimodal token-level annotations
Primary task
Schedule information extraction as token classification
Modalities
Image + text + 2D spatial layout
Coordinates… See the full description on the dataset page: https://huggingface.co/datasets/Ibnuck/poster-schedule-information-extraction.key_information_extractionreceipt_VLM_information_extractionhtml_to_json_information_extraction_dataset
HTML to JSON Information Extraction Dataset
Description
The html_to_json_information_extraction dataset is a collection of over 7300 HTML snippets and their extracted information in JSON.
These HTML have been sourced (scraped) from about 25 companies' career pages.
The dataset contains three splits - train, test, unseen_test.
This dataset has been built to fine tune SLMs & LLMs for the information extraction task.
train split
This split contains over 5700 pair… See the full description on the dataset page: https://huggingface.co/datasets/Jiraya/html_to_json_information_extraction_dataset.key-information-extraction
Neural Metrics · Straight at the task: pull the right fields out.
A key-information-extraction set aimed squarely at the core job - given a document, return the specific values that matter.
We use it for: benchmarking field-level precision and recall - measuring how often we hallucinate a plausible-but-absent value.
Attribution
This is an unmodified fork of nanonets/key_information_extraction, created by the Qwen team.
All weights, files and behaviour are… See the full description on the dataset page: https://huggingface.co/datasets/NeuralMetrics/key-information-extraction.kleister_nda_information_extraction
Kleister NDA — Information Extraction (orgrctera/kleister_nda_information_extraction)
Overview
This release packages the Kleister NDA split of the Kleister benchmark as rows suitable for information extraction (IE) evaluation. Each example points at a Non-Disclosure Agreement (NDA) document and specifies which attribute keys should be filled; the target is a JSON object of normalized string values for those keys (with null when a value is absent or not applicable).… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/kleister_nda_information_extraction.japanese-confidential-information-extraction-sft
Japanese Confidential Information Extraction — SFT Dataset
日本語テキストから社外秘の固有表現を抽出するタスク向けの SFT (Supervised Fine-Tuning) データセットです。
LFM2 系モデルの LoRA fine-tune を想定して構築されています。
タスク概要
入力テキスト(日本語)に含まれる機密情報を、11カテゴリの JSON として抽出します。
入力: 「山田太郎(yamada@example.co.jp)から請求書番号 INV-2024-0042 で
売上 ¥12,800,000 の見積書が届いた。」
出力: {
"address": [],
"company_name": [],
"email_address": ["yamada@example.co.jp"],
"human_name": ["山田太郎"],
"phone_number": [],
"account_identifier":… See the full description on the dataset page: https://huggingface.co/datasets/akiFQC/japanese-confidential-information-extraction-sft.sharedtask_nlpai4health_information_extraction_filteredkleister_charity_information_extraction
Kleister Charity — Information Extraction
Dataset description and background
Kleister Charity is one of two corpora introduced in Kleister: Key Information Extraction Datasets Involving Long Documents with Complex Layouts (ICDAR 2021). It targets Key Information Extraction (KIE) from long, formally structured English documents that mix scanned pages and born-digital PDFs, with complex visual layout (tables, multi-column text, headers, etc.).
The Charity split is built… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/kleister_charity_information_extraction.vital_articles_synthetic_information_extractionSynthetically annotated dataset for Wikipedia vital articles.
Relevant publication: https://arxiv.org/abs/2507.05997
pii_masking_300k_information_extraction
PII Masking 300k — Information Extraction
Dataset summary
This repository hosts a validation sample of the PII Masking 300k benchmark for the information extraction track: models must identify personally identifiable information (PII) in text and produce structured extractions (slot-filling JSON), optional token-level BIO labels, and span-based annotations for masking or redaction workflows.
The full PII Masking 300k suite is designed to stress-test privacy-preserving… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/pii_masking_300k_information_extraction.acled-information-extractioncustomer_service_information_extractioninformation-extraction-jsonreceipt_VLM_information_extractiondolly_information_extraction
Dataset Card for "dolly_information_extraction"
More Information needed
structured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetstructured-generation-information-extraction-vlms-openbmb-RLAIF-V-DatasetInformationExtractionQAstructured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetdolly-information-extraction
Dataset Card for "dolly-information-extraction"
More Information needed
structured-generation-information-extraction-vlms-openbmb-RLAIF-V-DatasetStructured_Information_Extractiondataset_information_extractionnanonets_key_information_extraction_masked_totalsalsama__Arabic-Information-Extraction-Corpus
