CoolFace
14 results

Doc Parsing

SoMarkAI /DocParsingBenchDocParsingBench is a document intelligence benchmark of 1,400 images, systematically collected and annotated from real business workflows. It is the first dataset to systematically catalogue the document elements most frequently encountered in enterprise settings, covering five major domains: finance, legal, scientific research, manufacturing, and education. 🆕 Latest Updates [2026.04.17] DocParsingBench evaluation toolkit release. Unified scoring is now available for the three… See the full description on the dataset page: https://huggingface.co/datasets/SoMarkAI/DocParsingBench.image1K<n<10K5 likes365 downloads5mo agoHugging Facekaihe /chinese_insurance_doc_parsing本数据集清洗自天池实验室公共数据集 结合原数据集的标注和pdf文档解析工具,构造了alpaca格式的数据: Instuction: 下列是直接从pdf原文件中提取出的某保险条款原文,pdf文件的字体排版存在一些空间结构,直接转换成字符串后会导致条款原文非常难以阅读。请把内容重新组织成清晰可读的格式。要求如下: 第一行是保险公司的全称 第二行是保险产品名 章节和子章节的序号统一用数字1-9表示 章节序号和章节名写在同一行,用空格进行间隔;章节具体内容放在下一行 章节和章节之间空一行 input: 使用pdfminer直接提取的字符串 中国太平洋人寿保险股份有限公司 个人税收递延型养老年金保险(2018 版) 产品基本条款 第一条 合同构成 个人税收递延型养老年金保险(2018 版)产品合同(以下简称“本合同”)由保险单及 所附个人税收递延型养老年金保险(2018 版)产品基本条款(以下简称“本合同基本条款 (2018 版)”)、个人税收递延型养老年金保险(2018 版)产品账户利益条款(以下简称“本 合同账户利益条款(2018… See the full description on the dataset page: https://huggingface.co/datasets/kaihe/chinese_insurance_doc_parsing.textn<1K10 likes59 downloads2y agoHugging FaceTravad98 /donut-docparsing-sogc-trademarks-1883-2001 Dataset Card for "donut-docparsing-sogc-trademarks-1883-2001" More Information needed text1K<n<10K0 likes21 downloads3y agoHugging Facetensorlake /OCRBenchV2-DocParsing-UpdatedGTOCRBenchV2-DocParsing-UpdatedGT is an improved ground truth for the document parsing subset of OCRBench V2, created and verified by Tensorlake. This version is used in Tensorlake’s OCR and document understanding benchmark comparisons. The dataset is intended for evaluation and research purposes only.For the original benchmark and other subsets, please refer to OCRBench V2 . texttext-generationn<1K0 likes16 downloads11mo agoHugging Face