CoolFace
Datasetpublic

cakradana-app/cakradana-kpu-filings-14k-pages

πŸ›οΈ Cakradana β€” KPU Campaign-Finance Filings Indonesian election campaign-finance filings, for document OCR and structured extraction πŸ“‹ Table of Contents 🎯 What this is πŸ“¦ What is in here πŸš€ Loading πŸ—‚οΈ Schema πŸ”’ Redaction ⚠️ Limitations πŸ“œ Licence and provenance 🎯 What this is Indonesian candidates and parties must file campaign-finance reports β€” LADK, LPSDK and LPPDK β€” with the KPU (Komisi Pemilihan Umum, the General Elections… See the full description on the dataset page: https://huggingface.co/datasets/cakradana-app/cakradana-kpu-filings-14k-pages.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes32downloads
Dataset Card

πŸ›οΈ Cakradana β€” KPU Campaign-Finance Filings

<div align="center">

<img src="assets/logo.png" alt="Cakradana" width="420">

Indonesian election campaign-finance filings, for document OCR and structured extraction

TrackAML 2.0 Winner ![Pages](#-what-is-in-here) ![Redacted](REDACTION.md) ![Website](https://cakradana.faizath.com)

</div>


πŸ“‹ Table of Contents


🎯 What this is

Indonesian candidates and parties must file campaign-finance reports β€” LADK, LPSDK and LPPDK β€” with the KPU (Komisi Pemilihan Umum, the General Elections Commission). They are published on regional KPU websites, and they are published almost entirely as scanned images with no text layer. A donation record that exists only as pixels cannot be audited at scale.

This dataset exists to close that gap. It supports two models:

modelinputtarget
document OCRpage imagepage text
extractionfiling plaintextstructured donation records

It was built for Cakradana, an AI system for transparency in Indonesian election financing, which placed 3rd in the Open Sector at the TrackAML 2.0 hackathon run by PPATK, Indonesia's Financial Transaction Reports and Analysis Center.


πŸ“¦ What is in here

14,540 pages from ~2,200 source documents across ~220 publishing hosts, as WebDataset shards plus gzipped JSONL manifests.

splitpagesshardswhat it is
train7,59521pages with a usable PDF text layer; label is pdftotext -layout output
eval_clean6882the same, from held-out hosts
eval_scan6,2578image-only pages; .txt is empty by design, labels live in teacher_labels
manifestrowswhat it is
extraction_pairs.jsonl.gz6,174(chunk text β†’ donation records) training pairs
teacher_labels.jsonl.gz6,257transcriptions for the eval_scan pages
pages.jsonl.gz14,540one row per published page: provenance and quality
shards.jsonl.gz31shard index with sha256

Splits are host-disjoint, not page-disjoint. Two filings from one kabupaten share a template, a font and one operator's scanning habits, so splitting below host level measures memorisation rather than generalisation.


πŸš€ Loading

python
from datasets import load_dataset

ocr = load_dataset("cakradana-app/cakradana-kpu-filings-14k-pages", "ocr")
# columns: __key__, jpg (PIL.Image), txt (str), json (dict)

pairs = load_dataset("cakradana-app/cakradana-kpu-filings-14k-pages", "extraction")
labels = load_dataset("cakradana-app/cakradana-kpu-filings-14k-pages", "teacher_labels")

eval_scan carries no .txt β€” that is what put those pages there. Join to the transcriptions on __key__:

python
truth = {row["key"]: row["transcription"] for row in labels["train"]}
for sample in ocr["eval_scan"]:
    reference = truth.get(sample["__key__"])

A runnable example is in `scripts/load_example.py`.


πŸ—‚οΈ Schema

OCR sample β€” three members per key, {doc_sha256[:16]}-p{page:04d}:

{key}.jpg grayscale render, long side 1280 px Β· {key}.txt the label Β· {key}.json provenance:

json
{"doc_sha256": "...", "page": 1, "source_url": "https://...pdf",
 "host": "kab-....kpu.go.id", "report_type": "LPSDK",
 "quality": {"extractor_agreement": 0.979, "digit_agreement": 1.0,
             "ink_fraction": 0.058, "char_count": 2240, "verdict": "clean"}}

`extraction_pairs` β€” unit, text, expected, is_hard_null, split, doc_sha256, chunk_index, start_page, end_page, donor_count, source_url, host, report_type, teacher_model.

expected is a JSON array of sparse donation records. Every field carries a *_span, and a span is the raw quoted substring, not an offset:

json
{"sender": "Hj. Nessy Ariyani", "amount": 750000, "date": "2015-10-15",
 "sender_span": "Hj. Nessy Ariyani", "amount_span": "750.000",
 "date_span": "15 OKTOBER 2015"}

The value is normalised; the span is what is printed on the page. A model that puts the normalised form in the span has that field rejected by the validator.

is_hard_null is exactly expected == []. 89% of chunks are hard nulls β€” covers, decision letters, expenditure tables. They are the fabrication test, not padding.

`teacher_labels` β€” key, model, transcription, records, n_returned, doc_sha256, source_url, host, report_type, page, host_split.

host_split is the host split used for extraction training. It is not the OCR bucket β€” every key in this file is an eval_scan page. Honour it when scoring, or your scan metrics leak.


πŸ”’ Redaction

3,333 of 17,873 pages (18.6%) were withheld because they carry β€” or sit beside something that carries β€” an identity number, phone number, address or date of birth.

Whole pages are removed rather than masked. The identifiers appear in the page images as well as the text, and the source PDFs were not retained, so the word boxes needed to blank a region are unrecoverable. Masking the text would have left the pixels intact.

Donor names are present. They are the extraction target, KPU publishes them, and a model trained on surrogate names would be trained on the wrong distribution.

[REDACTION.md](REDACTION.md) documents what was removed, how it was verified, and β€” importantly β€” what this process cannot guarantee. Read it before using this data about identifiable people.


⚠️ Limitations

  • β€”Publisher concentration. A small number of hosts contribute most donation records. Precision is measurable; generalisation to an unseen publisher is not.
  • β€”`train`/`eval_clean` labels are `pdftotext -layout` output, and that extractor is sometimes wrong. CER against them measures agreement with the extractor, not accuracy.
  • β€”`eval_scan` labels are model-generated. Read the trend, not the absolute.
  • β€”Column padding is normalised away in the published labels; line breaks are kept, because they carry table row structure.
  • β€”~22% duplicate renders were removed during the build. Every key now appears exactly once, which the verifier asserts.

πŸ“œ Licence and provenance

The underlying filings are public documents published by KPU. This repository redistributes derived renders and text under `LICENSE`; the filings themselves remain subject to Indonesian law on public information.

Every page keeps its source_url, host, report_type, page number and quality verdict, so any record can be traced to the document it came from.

Verify a copy of this dataset at any time:

bash
python verify_public.py --root .

<div align="center">

Made with ❀️ by the Cakradana Team · cakradana.faizath.com

</div>