document-review
document-review-data
Document Review Data
Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package.
Current Title Extraction Dataset Surface
Canonical prefix:
datasets/title_extraction/
Effective datasets:
datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/
datasets/title_extraction/evaluation/real_device_280_v1/
datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/
The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.document-review-source1k
New 1K title extraction corpus
Independent Task0178 expansion, parsed and machine annotated in Task0180, human-reviewed under Task0186. This is NOT the 1K subset sampled from the older 4K corpus. Private research use only; original document rights are not independently verified.
1000 unique source documents: Word, Excel, PDF and PPT each 250; English 500, simplified Chinese 400, traditional Chinese 100. All 1000 human reviews collected: 711 single-title, 194 multi-title, 54… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-source1k.apex-document-relevance-review
Apex Telecommunications Document Relevance Review (Synthetic)
A fully synthetic eDiscovery / internal-investigation dataset built to test whether an AI agent can
perform first-pass relevance coding on a realistic, adversarial document population.
The scenario
Apex Telecommunications, Inc. suspects Senior Sales Manager Michael Carter of sharing confidential
pricing and customer information with a competitor, Northstar Communications, before and during his
departure… See the full description on the dataset page: https://huggingface.co/datasets/Vrishab80/apex-document-relevance-review.Review_DocumentAI_Relevance_Document_Review
