CoolFace
Datasetpublic

mannycooper/document-review-data

Document Review Data Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package. Current Title Extraction Dataset Surface Canonical prefix: datasets/title_extraction/ Effective datasets: datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/ datasets/title_extraction/evaluation/real_device_280_v1/ datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/ The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.

sourceHugging Faceupdated 11d agoView on Hugging Face
3likes3kdownloads
Dataset Card

Document Review Data

Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package.

Current Title Extraction Dataset Surface

Canonical prefix:

text
datasets/title_extraction/

Effective datasets:

  • datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/
  • datasets/title_extraction/evaluation/real_device_280_v1/
  • datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/

The Dataset Viewer is intentionally configured to read only selected JSONL index/label files. The repository also contains XML, CSV, Office/PDF source files, metrics, and benchmark artifacts that should be consumed by path, not auto-inferred as viewer splits.

Legacy Review App Contract

The deployment contract for the older review app is:

text
manifests/hf_manifest.jsonl

Do not make this dataset public without clearing the source files.

mannycooper/document-review-data · CoolFace