CoolFace
20 results

information-extraction

Ibnuck /poster-schedule-information-extraction Multimodal Visual-Text Dataset for Poster Schedule Information Extraction A ready-to-train Indonesian document AI dataset combining pixels, OCR tokens, spatial layout, and BIO entity labels. Overview What it is 127 Indonesian seminar and religious-study event posters with multimodal token-level annotations Primary task Schedule information extraction as token classification Modalities Image + text + 2D spatial layout Coordinates… See the full description on the dataset page: https://huggingface.co/datasets/Ibnuck/poster-schedule-information-extraction.imagetoken-classificationn<1K1 likes112 downloads13d agoHugging Facenanonets /key_information_extractiontextquestion-answeringn<1K6 likes91 downloads1y agoHugging Faceamohseni /receipt_VLM_information_extractionimagen<1K2 likes52 downloads2y agoHugging FaceJiraya /html_to_json_information_extraction_dataset HTML to JSON Information Extraction Dataset Description The html_to_json_information_extraction dataset is a collection of over 7300 HTML snippets and their extracted information in JSON. These HTML have been sourced (scraped) from about 25 companies' career pages. The dataset contains three splits - train, test, unseen_test. This dataset has been built to fine tune SLMs & LLMs for the information extraction task. train split This split contains over 5700 pair… See the full description on the dataset page: https://huggingface.co/datasets/Jiraya/html_to_json_information_extraction_dataset.text1K<n<10K2 likes43 downloads1y agoHugging FaceNeuralMetrics /key-information-extraction Neural Metrics · Straight at the task: pull the right fields out. A key-information-extraction set aimed squarely at the core job - given a document, return the specific values that matter. We use it for: benchmarking field-level precision and recall - measuring how often we hallucinate a plausible-but-absent value. Attribution This is an unmodified fork of nanonets/key_information_extraction, created by the Qwen team. All weights, files and behaviour are… See the full description on the dataset page: https://huggingface.co/datasets/NeuralMetrics/key-information-extraction.textn<1K0 likes39 downloads1mo agoHugging Faceorgrctera /kleister_nda_information_extraction Kleister NDA — Information Extraction (orgrctera/kleister_nda_information_extraction) Overview This release packages the Kleister NDA split of the Kleister benchmark as rows suitable for information extraction (IE) evaluation. Each example points at a Non-Disclosure Agreement (NDA) document and specifies which attribute keys should be filled; the target is a JSON object of normalized string values for those keys (with null when a value is absent or not applicable).… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/kleister_nda_information_extraction.textn<1K0 likes36 downloads6mo agoHugging Face