CoolFace
Datasetpublic

amazon/ConfBench

FCC Invoices Verified Augmented Dataset Description FCC Invoices Verified Augmented is a document understanding benchmark dataset consisting of 75 real-world Federal Communications Commission (FCC) invoice documents, each augmented with up to 18 distinct document degradation pipelines. The dataset is designed to support confidence calibration research, OCR robustness evaluation, and key information extraction (KIE) under realistic noise conditions. Each document… See the full description on the dataset page: https://huggingface.co/datasets/amazon/ConfBench.

sourceHugging Facecc-by-nc-4.0updated 2mo agoView on Hugging Face
1likes1.5kdownloads
Dataset Card

FCC Invoices Verified Augmented

Dataset Description

FCC Invoices Verified Augmented is a document understanding benchmark dataset consisting of 75 real-world Federal Communications Commission (FCC) invoice documents, each augmented with up to 18 distinct document degradation pipelines. The dataset is designed to support confidence calibration research, OCR robustness evaluation, and key information extraction (KIE) under realistic noise conditions.

Each document has:

  • A clean original PDF
  • Up to 18 noisy versions generated by distinct Augraphy-based degradation pipelines
  • Verified ground-truth entity annotations

Total samples: 1,346 (75 documents × up to 18 noise pipelines)

Quick Start

Metadata + ground truth only (fastest, no PDFs):

python
from datasets import load_dataset

ds = load_dataset("amazon/ConfBench", data_files="data/test-00000-of-00001.parquet", split="train")

for sample in ds:
    print(sample["id"], sample["noise_variant"])
    print(sample["json_response"]["Agency"])

With PDFs (downloads the full dataset locally):

python
import pyarrow.parquet as pq
from huggingface_hub import snapshot_download
from pathlib import Path

local_dir = Path(snapshot_download(repo_id="amazon/ConfBench", repo_type="dataset"))

table = pq.read_table(local_dir / "data" / "test-00000-of-00001.parquet")
df = table.to_pandas()

# Access a PDF
pdf_path = local_dir / "pdfs" / df["id"][0]

# Access ground truth
gt = df["json_response"][0]
print(gt["Agency"], gt["GrossTotal"])

Or clone with git:

bash
git clone https://huggingface.co/datasets/amazon/ConfBench

Dataset Summary

PropertyValue
DomainLegal / Broadcast Advertising
Document TypeFCC Invoice (multi-page PDF)
# Base Documents75
# Noise Pipelinesup to 18 per document
# Total Samples1,346
Avg Pages per Doc~2
LicenseCC BY-NC 4.0

Dataset Structure

Repository Layout

amazon/ConfBench/
├── data/
│   └── test-00000-of-00001.parquet   # Metadata + ground truth (all variants)
├── pdfs/
│   ├── {doc_id}__original.pdf        # Original clean PDF
│   ├── {doc_id}__default.pdf         # Default noise variant
│   └── {doc_id}__{noise_variant}.pdf # Other noise variants

Ground Truth Schema (json_response)

json
{
  "document_class": {
    "type": "Invoice"
  },
  "split_document": {
    "page_indices": [0, 1]
  },
  "inference_result": {
    "Agency": "American Media & Advocacy Group",
    "Advertiser": "National Rifle Association",
    "GrossTotal": 15185.0,
    "PaymentTerms": "30 Days",
    "AgencyCommission": 2277.75,
    "NetAmountDue": 12907.25,
    "LineItems": [
      {
        "LineItemDescription": "M-F 1135p-1205a",
        "LineItemStartDate": "10/09/12",
        "LineItemEndDate": "10/09/12",
        "LineItemDays": "-T-----",
        "LineItemRate": 600.0
      }
    ]
  }
}

Noise Pipelines

18 distinct Augraphy-based degradation pipelines covering a wide range of real-world scan/print artifacts:

Augraphy Archetypes (pre-built pipelines)

PipelineDescription
defaultBalanced general-purpose degradation
archetype3Heavy post-processing effects
archetype4Minimal geometric distortions
archetype7Color and lighting variations
archetype9Texture-based degradations
archetype10Scanner artifact simulation
archetype11Complex multi-phase pipeline

Custom Pipelines (research-designed)

PipelineKey AugmentationsSimulates
custom12DirtyDrum + DirtyRollersScanner roller artifacts
custom13Stains + FoldingPhysical document damage
custom14BleedThrough + InkMottlingInk bleed and mottling
custom15Moire + ColorPaperScanning/aging effects
custom16ShadowCast + LightingGradientUneven lighting
custom17Jpeg + SubtleNoiseCompression artifacts
custom18Geometric + PageBorderAlignment issues
custom19BindingsAndFasteners + LetterpressBinding shadows
custom20Brightness + BadPhotoCopyPhotocopy quality
custom21WaterMark + NoisyLinesOverlaid artifacts
custom22Dithering + DotMatrixDot-matrix printing

Entity Fields

FieldTypeDescription
AgencystringAdvertising agency name
AdvertiserstringClient/advertiser name
GrossTotalfloatTotal invoice amount (USD)
PaymentTermsstringPayment terms (e.g., "30 Days")
AgencyCommissionfloatAgency commission amount (USD)
NetAmountDuefloatNet amount after commission (USD)
LineItemslistIndividual line items
LineItemDescriptionstringProgram/slot description
LineItemStartDatestringAiring start date (MM/DD/YY)
LineItemEndDatestringAiring end date (MM/DD/YY)
LineItemDaysstringDays of week pattern (e.g., "-T-----")
LineItemRatefloatCost per line item (USD)

Intended Use

This dataset is intended for:

  1. 1.Confidence calibration research — Measuring model confidence under varying degrees of document degradation.
  2. 2.OCR robustness evaluation — Benchmarking OCR and KIE systems on realistic noisy documents.
  3. 3.Document understanding — Evaluating models on structured information extraction from invoices (evaluation/benchmark use only; no training split is provided).
  4. 4.Noise impact analysis — Studying how specific noise types affect extraction accuracy per field.

Source Data

The 75 clean FCC invoices used in this dataset are taken from **RealKIE-FCC-Verified**, which re-annotated the FCC Invoices subset of the original **RealKIE** benchmark (Townsend et al., 2024) to fix two issues in the original annotations:

  1. 1.Line Item Grouping — fields previously treated as independent entries are grouped within each individual line item, aligning annotations with real-world invoice structure.
  2. 2.Annotation Corrections — erroneous values in the original annotations were corrected.

Starting from these 75 verified documents and their corrected ground truth (json_response), this dataset adds noise augmentation using the Augraphy library with OCR-safe parameter settings, producing up to 18 degraded variants per document for confidence calibration and OCR robustness research.


How to Cite This Dataset

If you use this dataset, please cite:

bibtex
@misc{islam2026confidencecalibration,
  title     = {FCC Invoices Verified Augmented: A Benchmark for Confidence Calibration in Document Understanding under Noise},
  author    = {Md Mofijul Islam and Mohammad Rostami and others},
  year      = {2026},
  note      = {Dataset available at Hugging Face},
  url       = {https://huggingface.co/datasets/amazon/ConfBench}
}

License

This dataset is released under CC BY-NC 4.0. The original FCC invoice documents are public records.