mannycooper/document-review-data
Document Review Data Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package. Current Title Extraction Dataset Surface Canonical prefix: datasets/title_extraction/ Effective datasets: datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/ datasets/title_extraction/evaluation/real_device_280_v1/ datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/ The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.
Document Review Data
Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package.
Current Title Extraction Dataset Surface
Canonical prefix:
datasets/title_extraction/Effective datasets:
datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/datasets/title_extraction/evaluation/real_device_280_v1/datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/
The Dataset Viewer is intentionally configured to read only selected JSONL index/label files. The repository also contains XML, CSV, Office/PDF source files, metrics, and benchmark artifacts that should be consumed by path, not auto-inferred as viewer splits.
Legacy Review App Contract
The deployment contract for the older review app is:
manifests/hf_manifest.jsonlDo not make this dataset public without clearing the source files.
