CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Ibnuck /poster-schedule-information-extraction Multimodal Visual-Text Dataset for Poster Schedule Information Extraction A ready-to-train Indonesian document AI dataset combining pixels, OCR tokens, spatial layout, and BIO entity labels. Overview What it is 127 Indonesian seminar and religious-study event posters with multimodal token-level annotations Primary task Schedule information extraction as token classification Modalities Image + text + 2D spatial layout Coordinates… See the full description on the dataset page: https://huggingface.co/datasets/Ibnuck/poster-schedule-information-extraction.imagetoken-classificationn<1K1 likes112 downloads13d agoHugging Face02nanonets /key_information_extractiontextquestion-answeringn<1K6 likes91 downloads1y agoHugging Face03amohseni /receipt_VLM_information_extractionimagen<1K2 likes52 downloads2y agoHugging Face04Jiraya /html_to_json_information_extraction_dataset HTML to JSON Information Extraction Dataset Description The html_to_json_information_extraction dataset is a collection of over 7300 HTML snippets and their extracted information in JSON. These HTML have been sourced (scraped) from about 25 companies' career pages. The dataset contains three splits - train, test, unseen_test. This dataset has been built to fine tune SLMs & LLMs for the information extraction task. train split This split contains over 5700 pair… See the full description on the dataset page: https://huggingface.co/datasets/Jiraya/html_to_json_information_extraction_dataset.text1K<n<10K2 likes43 downloads1y agoHugging Face05NeuralMetrics /key-information-extraction Neural Metrics · Straight at the task: pull the right fields out. A key-information-extraction set aimed squarely at the core job - given a document, return the specific values that matter. We use it for: benchmarking field-level precision and recall - measuring how often we hallucinate a plausible-but-absent value. Attribution This is an unmodified fork of nanonets/key_information_extraction, created by the Qwen team. All weights, files and behaviour are… See the full description on the dataset page: https://huggingface.co/datasets/NeuralMetrics/key-information-extraction.textn<1K0 likes39 downloads1mo agoHugging Face06orgrctera /kleister_nda_information_extraction Kleister NDA — Information Extraction (orgrctera/kleister_nda_information_extraction) Overview This release packages the Kleister NDA split of the Kleister benchmark as rows suitable for information extraction (IE) evaluation. Each example points at a Non-Disclosure Agreement (NDA) document and specifies which attribute keys should be filled; the target is a JSON object of normalized string values for those keys (with null when a value is absent or not applicable).… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/kleister_nda_information_extraction.textn<1K0 likes36 downloads6mo agoHugging Face07akiFQC /japanese-confidential-information-extraction-sft Japanese Confidential Information Extraction — SFT Dataset 日本語テキストから社外秘の固有表現を抽出するタスク向けの SFT (Supervised Fine-Tuning) データセットです。 LFM2 系モデルの LoRA fine-tune を想定して構築されています。 タスク概要 入力テキスト(日本語)に含まれる機密情報を、11カテゴリの JSON として抽出します。 入力: 「山田太郎(yamada@example.co.jp)から請求書番号 INV-2024-0042 で 売上 ¥12,800,000 の見積書が届いた。」 出力: { "address": [], "company_name": [], "email_address": ["yamada@example.co.jp"], "human_name": ["山田太郎"], "phone_number": [], "account_identifier":… See the full description on the dataset page: https://huggingface.co/datasets/akiFQC/japanese-confidential-information-extraction-sft.texttext-generation10K<n<100K2 likes33 downloads4mo agoHugging Face08vinaybabu /sharedtask_nlpai4health_information_extraction_filteredtext10K<n<100K0 likes25 downloads11mo agoHugging Face09orgrctera /kleister_charity_information_extraction Kleister Charity — Information Extraction Dataset description and background Kleister Charity is one of two corpora introduced in Kleister: Key Information Extraction Datasets Involving Long Documents with Complex Layouts (ICDAR 2021). It targets Key Information Extraction (KIE) from long, formally structured English documents that mix scanned pages and born-digital PDFs, with complex visual layout (tables, multi-column text, headers, etc.). The Charity split is built… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/kleister_charity_information_extraction.textother1K<n<10K0 likes25 downloads6mo agoHugging Face10nicpopovic /vital_articles_synthetic_information_extractionSynthetically annotated dataset for Wikipedia vital articles. Relevant publication: https://arxiv.org/abs/2507.05997 text1K<n<10K1 likes23 downloads1y agoHugging Face11orgrctera /pii_masking_300k_information_extraction PII Masking 300k — Information Extraction Dataset summary This repository hosts a validation sample of the PII Masking 300k benchmark for the information extraction track: models must identify personally identifiable information (PII) in text and produce structured extractions (slot-filling JSON), optional token-level BIO labels, and span-based annotations for masking or redaction workflows. The full PII Masking 300k suite is designed to stress-test privacy-preserving… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/pii_masking_300k_information_extraction.texttoken-classificationn<1K0 likes22 downloads6mo agoHugging Face12vinaykudari /acled-information-extractiontext10K<n<100K3 likes21 downloads4y agoHugging Face13jonathansuru /customer_service_information_extractiontextfeature-extractionn<1K4 likes19 downloads3y agoHugging Face14kamizane /information-extraction-jsontextn<1K0 likes17 downloads7mo agoHugging Face15Ryann5200 /receipt_VLM_information_extractionimagen<1K0 likes14 downloads1mo agoHugging Face16lionelchg /dolly_information_extraction Dataset Card for "dolly_information_extraction" More Information needed text1K<n<10K4 likes13 downloads3y agoHugging Face17davidberenstein1957 /structured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetimagen<1K0 likes13 downloads2y agoHugging Face18Hoang123 /structured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetimagen<1K0 likes12 downloads1y agoHugging Face19OnnieNLP /InformationExtractionQAtextquestion-answeringn<1K0 likes12 downloads1y agoHugging Face20paulwoodward /structured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetimagen<1K0 likes10 downloads2y agoHugging Face21lionelchg /dolly-information-extraction Dataset Card for "dolly-information-extraction" More Information needed text1K<n<10K0 likes9 downloads3y agoHugging Face22chandc /structured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetimagen<1K0 likes8 downloads1y agoHugging Face23IoakeimE /Structured_Information_Extractiontext100K<n<1M0 likes7 downloads5mo agoHugging Face24Maeda-miyazaki /dataset_information_extractiontextn<1K0 likes6 downloads3y agoHugging Face25Hushan-10 /nanonets_key_information_extraction_masked_totalimagen<1K0 likes6 downloads8mo agoHugging Face26alielfilali01 /salsama__Arabic-Information-Extraction-Corpus0 likes3 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.