datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
uds-governance-receipts
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
UDS Governance Receipts — Decision Audit Log
Doctrine v11 LOCKED. No marketing. Every number resolves to a CI log, a Lean proof, or a Zenodo DOI.
Append-only log of DSSE-signed governance decision receipts for the Unified Deployment Substrate (UDS) mesh. Each record captures:… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/uds-governance-receipts.synthetic-receipts-ocr
synthetic-receipts-ocr
32,000 synthetic thermal receipts across 5 locales (US/UK/DE/IT/FR) — each a clean render plus a photo-degraded twin, with pixel-exact word boxes, full transcription, and structured KIE fields.
Samples train-000357 (US), train-000073 (UK), eval-001179 (DE), train-000222 (IT), train-000711 (FR) — real dataset rows, not mockups. Each receipt is its sample's image_photo, cut out along its own homography quad; no retouching beyond composition.
Built for… See the full description on the dataset page: https://huggingface.co/datasets/albertobarnabo/synthetic-receipts-ocr.uds-spans-receipts
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
UDS Spans Receipts — OTel Governance Audit Log
Doctrine v11 LOCKED. No marketing. Every number resolves to a CI log, a Lean proof, or a Zenodo DOI.
Append-only audit log of DSSE-signed OpenTelemetry spans emitted by the UDS mesh governance layer. Each span record includes: operation… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/uds-spans-receipts.invoices-and-receipts_ocr_v1
Dataset Card for "invoices-and-receipts_ocr_v1"
More Information needed
governed-receipts-bench
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
Governed Receipts Bench · a conformance corpus for the governed-receipt spec
A small benchmark corpus of governance decision receipts for the open
governed-receipt-spec.
bench.jsonl declares one expected outcome for each case under the spec's
dependency-free offline verifier. With the… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/governed-receipts-bench.invoices-and-receipts_ocr_v2
Dataset Card for "invoices-and-receipts_ocr_v2"
Usage
from datasets import load_dataset
dataset = load_dataset("mychen76/invoices-and-receipts_ocr_v2")
dataset
More Information needed
receipts-google-ocrreceipts-finetune-v2receipts-finetune-v3nano-receipts
🧾 Nano Receipts Dataset
A diverse collection of 2428 hyper-realistic synthetic receipt images generated using state-of-the-art text-to-image AI models.
🚀 Quick Start
from datasets import load_dataset
# Load dataset (fast parquet format!)
dataset = load_dataset("34data/nano-receipts")
# Access images
image = dataset["train"][0]["image"] # PIL Image
filename = dataset["train"][0]["filename"]
📊 Dataset Details
Total Images: 2428 receipts
Format:… See the full description on the dataset page: https://huggingface.co/datasets/34data/nano-receipts.receipts-dataset-v1
Dataset Card for "receipts-dataset-v1"
More Information needed
receipts-i2ireceipts-finetune-v1ocr-receipts-text-detectionThe Grocery Store Receipts Dataset is a collection of photos captured from various
**grocery store receipts**. This dataset is specifically designed for tasks related to
**Optical Character Recognition (OCR)** and is useful for retail.
Each image in the dataset is accompanied by bounding box annotations, indicating the
precise locations of specific text segments on the receipts. The text segments are
categorized into four classes: **item, store, date_time and total**.donut_receiptsreceipts-ocr-dataset
Receipt OCR Dataset
A dataset of receipt photos with structured JSON extraction labels for fine-tuning vision-language models on document OCR tasks.
Dataset
230 receipt images labeled with structured JSON extracted via Gemini, covering a variety of merchants, formats, and receipt layouts.
Format
Column
Type
Description
image
Image
Receipt photo (JPEG)
text
string
Extracted receipt data as JSON
JSON Schema
{
"merchantName":… See the full description on the dataset page: https://huggingface.co/datasets/docjay131/receipts-ocr-dataset.nano-receipts
🧾 Nano Receipts Dataset
A diverse collection of 2428 hyper-realistic synthetic receipt images generated using state-of-the-art text-to-image AI models.
🚀 Quick Start
from datasets import load_dataset
# Load dataset (fast parquet format!)
dataset = load_dataset("34data/nano-receipts")
# Access images
image = dataset["train"][0]["image"] # PIL Image
filename = dataset["train"][0]["filename"]
📊 Dataset Details
Total Images: 2428 receipts… See the full description on the dataset page: https://huggingface.co/datasets/samarth010/nano-receipts.ds_receipts_v2_train
Dataset Card for "ds_receipts_v2_train"
More Information needed
taliesin-receipts
Taliesin — Verification Receipts
The cryptographic proof behind every public claim Corbenic AI makes about Taliesin, our lossless
external-memory engine. Don't trust us — check the hashes. Every result here is a SHA-256 receipt or
a structured NDJSON/JSON record produced by the actual tests.
The Taliesin engine itself is proprietary and is not in this repository. What is here is the
evidence: the hashes and structured outputs that let you verify the claims without our software.… See the full description on the dataset page: https://huggingface.co/datasets/Corbenic/taliesin-receipts.Invoice_and_receiptsassay-receipts
Assay Receipt Corpus
Signed internal receipts from an inference provider that is sometimes cheating, together
with the verdict an auditor reached on each one and the ground truth of which model actually
served the request.
Each row is a real receipt, not a summary statistic: it carries the prompt and output token
ids, the JL-projected sketch of the provider's hidden_states, the sign/rank invariants, and
an HMAC signature. With the gpt2 weights you can recompute the sketch… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/assay-receipts.receipts-v1synth_receipts_ocrus_federal_debt_and_receipts
بدهی و کسری بودجهٔ دولت فدرال آمریکا
بدهی کل دولت فدرال آمریکا بهصورت روزانه، و دریافتی و پرداختی ماهانهٔ خزانهداری. مبنای سنجش سرعت رشد کسری — عددی که هزینهٔ استقراض جهانی را تعیین میکند.
پوشش: 1372-01-12 → 1405-06-30 · تناوب: روزانه (بدهی) و ماهانه (کسری) · سطح: ایالات متحده
تعداد مشاهده: 8,673 · تعداد مکان: 1
منبع: خزانهداری ایالات متحده — Fiscal Data — https://fiscaldata.treasury.gov
شاخصها
شناسه
نام
واحد
debt.us_federal_total
بدهی کل… See the full description on the dataset page: https://huggingface.co/datasets/Farmaanaa/us_federal_debt_and_receipts.scbe-governance-receipts-v1
Status: experimental. Experiment-specific slice. Primary public dataset: scbe-aethermoore-training-data.
SCBE Governance Receipts v1
Schema: scbe_governed_dataset_v1
Receipt schema: scbe_governance_receipt_v1
Built: 2026-05-04T23:45:37Z
Rows: 40
Dataset ID: scbe-governance-receipts-v1
What this is
A governed dataset where every row carries a full 34-field SCBE
governance receipt (poly-embedded JEPA fingerprint + tri-vector
cross-braid hash + Sacred Egg ring seal… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-governance-receipts-v1.invoices-and-receipts_ocr_v1
Dataset Card for "invoices-and-receipts_ocr_v1"
More Information needed
receipts-agent-claims
Receipts — Agent Claim Transcripts
Every transcript from the Receipts benchmark — one row per trial, graded by a pytest exit code rather than by another model.
424 runs on claude-haiku-4-5, plus 6 pilot runs on gemini-2.5-flash via aider. All trials are committed. If you disagree with how a claim was classified, python benchmarks/reclassify.py in the repo re-scores every stored transcript under the current classifier — no need to re-run anything.
What the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Hachiman94/receipts-agent-claims.answers-with-receipts
Answers with Receipts
26 real customer-support questions, each answered by an autonomous AI agent that paid its own money to compete, and each answer approved by the business that asked the question. Every row carries the on-chain transaction that paid the agent.
The preference label in this dataset is backed by a payment, not a click.
Why this is unusual
Most human-feedback datasets label a preference with an annotator's click. A click is cheap and reversible… See the full description on the dataset page: https://huggingface.co/datasets/deskcrew/answers-with-receipts.invoices-and-receipts_ocr_v1
Dataset Card for "invoices-and-receipts_ocr_v1"
More Information needed
invoices-and-receipts_ocr_cleaned
invoices-and-receipts_ocr_cleaned
The invoices-and-receipts_ocr__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
2,221
QA turns
7,696
answers rewritten by the cleaning pass
1,517
QA created by the cleaning pass (new_qa)
4,651 (60.4%)
shards
4
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/invoices-and-receipts_ocr_cleaned.
