datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
InvoiceBenchmark
InvoiceBenchmark
200 synthetic invoices with cent-perfect ground truth, designed to measure the one thing language models are supposed to be able to do: read a number.
The Pitch
Invoice processing is the use case every enterprise AI pitch deck opens with. The numbers are either right or wrong, and the distance between right and wrong can be measured to the cent. This dataset exists because we ran the experiment and discovered that the gap between "this looks easy" and… See the full description on the dataset page: https://huggingface.co/datasets/jngb-labs/InvoiceBenchmark.Invoice-to-Json
Invoice-to-Json Dataset
Dataset Description
Dataset Summary
Invoice-to-Json is a dataset designed for document understanding and information extraction tasks. It consists of document images paired with questions and answers, specifically focused on extracting structured information (JSON format) from documents.
Supported Tasks
Document Question Answering: The dataset supports training models to answer questions about document content
Information… See the full description on the dataset page: https://huggingface.co/datasets/shubh303/Invoice-to-Json.maritime-demurrage-detention-invoice-event-coherence-risk-v0.1What this repo is for
Catch billing that does not match the container event record.
You use it to flag
charges while customs or terminal holds block pickup
charges when appointment scarcity blocks pickup
weak billing when free time terms are unclear
long return cycles that drive cost spikes
Why it matters
This is daily friction for shippers and forwarders.
It drives disputes, cashflow drag, and churn.
Invoice-to-Json
Invoice-to-Json Dataset
Dataset Description
Dataset Summary
Invoice-to-Json is a dataset designed for document understanding and information extraction tasks. It consists of document images paired with questions and answers, specifically focused on extracting structured information (JSON format) from documents.
Supported Tasks
Document Question Answering: The dataset supports training models to answer questions about document content
Information… See the full description on the dataset page: https://huggingface.co/datasets/acoustichao/Invoice-to-Json.carbon-mrv-invoice-emissions
Synthetic Carbon MRV Invoice-to-Emissions Dataset
A synthetic dataset that models the core pipeline used by carbon Measurement,
Reporting & Verification (MRV) platforms: turning a business document line
item (invoice, fuel receipt, electricity bill, freight charge) into a
GHG Protocol Scope 1 / 2 / 3 classification and a calculated emissions
value.
It was built as reference dataset for learning and prototyping —
specifically for training/evaluating models that do:
Scope… See the full description on the dataset page: https://huggingface.co/datasets/Samarth-27/carbon-mrv-invoice-emissions.INVOICE_ANNOTATION_V1invoice_datasetzephyr-7b-beta-invoices
Zephyr-7B-Beta Customer Support Chatbot
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Introduction
Welcome to the zephyr-7b-beta-invoices repository! This project leverages the Zephyr-7B-Beta model trained on the "Bitext-Customer-Support-LLM-Chatbot-Training-Dataset" to create a state-of-the-art customer support chatbot. Our goal is to provide an efficient and accurate chatbot for handling invoice-related… See the full description on the dataset page: https://huggingface.co/datasets/erfanvaredi/zephyr-7b-beta-invoices.INVOICE_ANNOTATION_V2invoices_v2invoiceInvoice_datainvoiceinvoice-ocr-datasetInvoicesinvoices_v2
