datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
clearocr-invoice-document-ai
clearOCR Invoice Document AI Dataset
This dataset shows a complete invoice document AI workflow built around clearOCR.
It contains 423 high-confidence invoice examples with:
original invoice images,
OCR text generated by clearOCR,
Markdown reconstruction of the document,
structured invoice JSON generated by a local fine-tuned extraction model,
visual verification metadata.
The dataset demonstrates how clearOCR can serve as the OCR layer in an invoice automation pipeline where… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/clearocr-invoice-document-ai.ru-invoice-extraction-benchmark
Набор для извлечения данных из русскоязычных счетов
50 синтетических русскоязычных счетов с эталонными ответами в JSON. Набор не привязан к определённой модели и предназначен для проверки систем извлечения структурированных данных из текста.
Разделы
Раздел
Документы
Назначение
development
30
разработка шаблонов и примеров
validation
10
выбор настроек
test
10
итоговая оценка
Раздел test зафиксирован для версии 1. Его нельзя использовать для… See the full description on the dataset page: https://huggingface.co/datasets/necrasov-ilya/ru-invoice-extraction-benchmark.invoice-extraction-dataset-v2
📑 Invoice Extraction Dataset v2
This dataset is specifically designed for fine-tuning Large Language Models (LLMs) to perform Structured Data Extraction.
It contains 601 unique examples of messy, unstructured invoice and receipt descriptions paired with their clean, machine-readable JSON counterparts.
🚀 Dataset Purpose
General LLMs often include conversational filler when asked for JSON. This dataset was built to train models to:
Ignore "Noise": Ignore conversational… See the full description on the dataset page: https://huggingface.co/datasets/manuelaschrittwieser/invoice-extraction-dataset-v2.japanese-invoice-receipt-extraction-eval
証憑 · Shōhyō — 日本語 証憑(請求書・領収書・支払通知書)構造化抽出ベンチマーク(無料サンプル)
証憑(しょうひょう)=取引の事実を証明する書類(請求書・領収書など)の会計用語。
「文書テキスト → JSON 抽出」パイプラインの精度を測るための 正解付き評価データセット の無料サンプルです。
config
文書タイプ
無料サンプル
invoice
請求書
20 件
receipt
領収書
10 件
payment_notice
支払通知書 / 仕入明細書
15 件
いずれもインボイス制度(適格請求書等保存方式)に対応。
既存のHF日本語帳票データはOCR・画像系が中心です。本データは 画像でなくテキスト→JSON を対象にし、和暦・軽減税率・源泉徴収・収入印紙税・相殺控除のロジックを正解側で算術検証してあります。
実在の企業名・個人情報は含みません(すべて合成)。本サンプルは無料・評価/検証用途で配布します。
このサンプルの位置づけ(凍結版)… See the full description on the dataset page: https://huggingface.co/datasets/Aulvem/japanese-invoice-receipt-extraction-eval.invoice_dataset
MindMap Enterprise Invoice Extraction & Fine-Tuning Dataset
Production-grade benchmark and fine-tuning dataset built from 1,007 enterprise PDF invoices for extracting structured JSON metadata matching strict target Pydantic schemas.
Dataset Structure & Splits
train (1,074 examples): 80% training split containing English and Hindi Devanagari invoice document pairs.
validation (134 examples): 10% validation split.
test_golden (135 examples): 10% held-out test split… See the full description on the dataset page: https://huggingface.co/datasets/Msduck/invoice_dataset.zephyr-7b-beta-invoices
Zephyr-7B-Beta Customer Support Chatbot
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Introduction
Welcome to the zephyr-7b-beta-invoices repository! This project leverages the Zephyr-7B-Beta model trained on the "Bitext-Customer-Support-LLM-Chatbot-Training-Dataset" to create a state-of-the-art customer support chatbot. Our goal is to provide an efficient and accurate chatbot for handling invoice-related… See the full description on the dataset page: https://huggingface.co/datasets/erfanvaredi/zephyr-7b-beta-invoices.
