amazon/ConfBench
FCC Invoices Verified Augmented Dataset Description FCC Invoices Verified Augmented is a document understanding benchmark dataset consisting of 75 real-world Federal Communications Commission (FCC) invoice documents, each augmented with up to 18 distinct document degradation pipelines. The dataset is designed to support confidence calibration research, OCR robustness evaluation, and key information extraction (KIE) under realistic noise conditions. Each document… See the full description on the dataset page: https://huggingface.co/datasets/amazon/ConfBench.
FCC Invoices Verified Augmented
Dataset Description
FCC Invoices Verified Augmented is a document understanding benchmark dataset consisting of 75 real-world Federal Communications Commission (FCC) invoice documents, each augmented with up to 18 distinct document degradation pipelines. The dataset is designed to support confidence calibration research, OCR robustness evaluation, and key information extraction (KIE) under realistic noise conditions.
Each document has:
- A clean original PDF
- Up to 18 noisy versions generated by distinct Augraphy-based degradation pipelines
- Verified ground-truth entity annotations
Total samples: 1,346 (75 documents × up to 18 noise pipelines)
Quick Start
Metadata + ground truth only (fastest, no PDFs):
from datasets import load_dataset
ds = load_dataset("amazon/ConfBench", data_files="data/test-00000-of-00001.parquet", split="train")
for sample in ds:
print(sample["id"], sample["noise_variant"])
print(sample["json_response"]["Agency"])With PDFs (downloads the full dataset locally):
import pyarrow.parquet as pq
from huggingface_hub import snapshot_download
from pathlib import Path
local_dir = Path(snapshot_download(repo_id="amazon/ConfBench", repo_type="dataset"))
table = pq.read_table(local_dir / "data" / "test-00000-of-00001.parquet")
df = table.to_pandas()
# Access a PDF
pdf_path = local_dir / "pdfs" / df["id"][0]
# Access ground truth
gt = df["json_response"][0]
print(gt["Agency"], gt["GrossTotal"])Or clone with git:
git clone https://huggingface.co/datasets/amazon/ConfBenchDataset Summary
Dataset Structure
Repository Layout
amazon/ConfBench/
├── data/
│ └── test-00000-of-00001.parquet # Metadata + ground truth (all variants)
├── pdfs/
│ ├── {doc_id}__original.pdf # Original clean PDF
│ ├── {doc_id}__default.pdf # Default noise variant
│ └── {doc_id}__{noise_variant}.pdf # Other noise variants
Ground Truth Schema (json_response)
{
"document_class": {
"type": "Invoice"
},
"split_document": {
"page_indices": [0, 1]
},
"inference_result": {
"Agency": "American Media & Advocacy Group",
"Advertiser": "National Rifle Association",
"GrossTotal": 15185.0,
"PaymentTerms": "30 Days",
"AgencyCommission": 2277.75,
"NetAmountDue": 12907.25,
"LineItems": [
{
"LineItemDescription": "M-F 1135p-1205a",
"LineItemStartDate": "10/09/12",
"LineItemEndDate": "10/09/12",
"LineItemDays": "-T-----",
"LineItemRate": 600.0
}
]
}
}Noise Pipelines
18 distinct Augraphy-based degradation pipelines covering a wide range of real-world scan/print artifacts:
Augraphy Archetypes (pre-built pipelines)
Custom Pipelines (research-designed)
Entity Fields
Intended Use
This dataset is intended for:
- Confidence calibration research — Measuring model confidence under varying degrees of document degradation.
- OCR robustness evaluation — Benchmarking OCR and KIE systems on realistic noisy documents.
- Document understanding — Evaluating models on structured information extraction from invoices (evaluation/benchmark use only; no training split is provided).
- Noise impact analysis — Studying how specific noise types affect extraction accuracy per field.
Source Data
The 75 clean FCC invoices used in this dataset are taken from **RealKIE-FCC-Verified**, which re-annotated the FCC Invoices subset of the original **RealKIE** benchmark (Townsend et al., 2024) to fix two issues in the original annotations:
- Line Item Grouping — fields previously treated as independent entries are grouped within each individual line item, aligning annotations with real-world invoice structure.
- Annotation Corrections — erroneous values in the original annotations were corrected.
Starting from these 75 verified documents and their corrected ground truth (json_response), this dataset adds noise augmentation using the Augraphy library with OCR-safe parameter settings, producing up to 18 degraded variants per document for confidence calibration and OCR robustness research.
How to Cite This Dataset
If you use this dataset, please cite:
@misc{islam2026confidencecalibration,
title = {FCC Invoices Verified Augmented: A Benchmark for Confidence Calibration in Document Understanding under Noise},
author = {Md Mofijul Islam and Mohammad Rostami and others},
year = {2026},
note = {Dataset available at Hugging Face},
url = {https://huggingface.co/datasets/amazon/ConfBench}
}License
This dataset is released under CC BY-NC 4.0. The original FCC invoice documents are public records.
