CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01D2I-CUHK-Shenzhen /FormStruct-Bench FormStruct-Bench Dataset Description FormStruct-Bench is a multilingual benchmark for extracting the semantic and spatial structure of forms from document images. The repository combines a 7,000-page main benchmark, a controlled visual-degradation set, and template-level layout annotations. It supports evaluation of vision-language models and document AI systems on hierarchical key-value extraction, document structure recovery, region localization, table and… See the full description on the dataset page: https://huggingface.co/datasets/D2I-CUHK-Shenzhen/FormStruct-Bench.imageimage-to-text1K<n<10K1 likes8.6k downloads2mo agoHugging Face02form-data-experiments /ANC-formsimage10K<n<100K0 likes189 downloads3y agoHugging Face03tanmaykm /indian_dance_formsThis dataset is taken from https://www.kaggle.com/datasets/aditya48/indian-dance-form-classification but is originally from the Hackerearth deep learning contest of identifying Indian dance forms. All the credits of dataset goes to them. Content The dataset consists of 599 images belonging to 8 categories, namely manipuri, bharatanatyam, odissi, kathakali, kathak, sattriya, kuchipudi, and mohiniyattam. The original dataset was quite unstructured and all the images were put together.… See the full description on the dataset page: https://huggingface.co/datasets/tanmaykm/indian_dance_forms.imageimage-classification1K<n<10K1 likes157 downloads3y agoHugging Face04Symage /synthetic-us-forms-preview SymageDocs — Synthetic US Forms Preview A small, CC-BY-4.0, fully synthetic document-AI training set: 525 labeled page images across six families of US business and government forms, each page shipping FUNSD ground truth plus a LayoutLM-ready token/bbox/tag view. This is a preview subset. It exists so you can load real output from the SymageDocs generator, inspect the label quality, and decide whether generating your own corpus is worth your time — without an account, an email… See the full description on the dataset page: https://huggingface.co/datasets/Symage/synthetic-us-forms-preview.imagetoken-classificationn<1K2 likes110 downloads1mo agoHugging Face05ift /handwriting_forms Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/ift/handwriting_forms.imagefeature-extraction1K<n<10K14 likes78 downloads3y agoHugging Face06hyturing /US_tax_forms_donut NIST-SD2 (US Tax Forms) Donut Dataset Processed version of NIST Special Database 2 for document understanding tasks, formatted for use with the Donut architecture. Contains 5,590 annotated document images (5,031 train, 559 test) across 20 tax form classes. Description Curated by: National Institute of Standards and Technology (NIST) License: MIT Total Size: 947.33 MB Annotations: Class labels (20 tax form types) Full text ground truth image1K<n<10K4 likes54 downloads2y agoHugging Face07Victorgonl /UFLA-FORMS UFLA-FORMS: an Academic Forms Dataset for Information Extraction in the Portuguese Language About UFLA-FORMS is a manually labeled dataset of document forms in Brazilian Portuguese extracted from the domains of the Federal University of Lavras (UFLA). The dataset emphasizes the hierarchical structure between the entities of a document through their relationships, in addition to the extraction of key-value pairs. Samples were labeled using ToolRI. Overview… See the full description on the dataset page: https://huggingface.co/datasets/Victorgonl/UFLA-FORMS.imagetoken-classificationn<1K1 likes49 downloads2y agoHugging Face08Elliot-Data /handwriting_forms_cleanedgated handwriting_forms_cleaned The handwriting_forms__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning. images 1,360 QA turns 5,198 answers rewritten by the cleaning pass 166 QA created by the cleaning pass (new_qa) 3,860 (74.3%) shards 1 How this was cleaned A vision-language model read each image together with its QA and judged the item. The pass is not a filter that only removes rows — it rewrites answers it finds… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/handwriting_forms_cleaned.imagevisual-question-answering1K<n<10K0 likes25 downloads12d agoHugging Face09saurabh1896 /OMR-forms Dataset Card for "OMR-forms" More Information needed imagen<1K1 likes23 downloads3y agoHugging Face10Ronysalem /medical-forms-datasetimagen<1K0 likes22 downloads2y agoHugging Face11elliot-mllm /handwriting_forms_cleanedgated handwriting_forms_cleaned The handwriting_forms__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning. images 1,360 QA turns 5,198 answers rewritten by the cleaning pass 166 QA created by the cleaning pass (new_qa) 3,860 (74.3%) shards 1 How this was cleaned A vision-language model read each image together with its QA and judged the item. The pass is not a filter that only removes rows — it rewrites answers it finds… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/handwriting_forms_cleaned.imagevisual-question-answering1K<n<10K0 likes22 downloads22d agoHugging Face12Symage /coherent-forms-1040-cms1500-i9gated SymageDocs — Coherent US Tax / Health / Employment Forms (FUNSD) A fully synthetic document-AI training set: three US forms — IRS Form 1040, CMS-1500, and USCIS Form I-9 — filled from the same synthetic identity, so name / SSN / address / employer flow consistently across all three renderings. Each page ships with FUNSD ground truth (word boxes, entity labels, key–value linking) plus a LayoutLMv3-ready token/bbox/tag view. 3,000 page-level image + annotation rows (train 2,400 /… See the full description on the dataset page: https://huggingface.co/datasets/Symage/coherent-forms-1040-cms1500-i9.imagetoken-classification1K<n<10K4 likes21 downloads3mo agoHugging Face13SiddharthStarr /indian_dance_formsThis dataset is taken from https://www.kaggle.com/datasets/aditya48/indian-dance-form-classification but is originally from the Hackerearth deep learning contest of identifying Indian dance forms. All the credits of dataset goes to them. Content The dataset consists of 599 images belonging to 8 categories, namely manipuri, bharatanatyam, odissi, kathakali, kathak, sattriya, kuchipudi, and mohiniyattam. The original dataset was quite unstructured and all the images were put together.… See the full description on the dataset page: https://huggingface.co/datasets/SiddharthStarr/indian_dance_forms.imageimage-classification1K<n<10K1 likes10 downloads6mo agoHugging Face14nnul /forms-from-rvl-cdipimage10K<n<100K0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.