ai4data/data-snapshot
Dataset card for data-snapshot This dataset was introduced in the paper Benchmarking Open-Source Layout Detection Models for Data Snapshot Extraction from Institutional Documents. The source code for the benchmark and dataset extraction is available on GitHub: worldbank/ai4data. Dataset summary The data-snapshot dataset is an annotated corpus designed for the evaluation and development of models for extracting data snapshots from PDF documents. A data snapshot is… See the full description on the dataset page: https://huggingface.co/datasets/ai4data/data-snapshot.
Minor edit to force viewer to reset
Update readme
Update citation
Update license info
Add paper link, code link, and citation information (#2)
Change username to ai4data
Add loading dataset examples
Update docs
Fix metadata
Fix prwp and refugee metadata
Fix unhcr metadata
Fix refugee metadata
Fix prwp metadata
Fix unhcr metadata
Fix unhcr metadata
Add metadata files
Update schema file
Fix timestamp issue
Upload all annotations
Add documents
Add snapshots
Update subsets
Update unhcr metadata
Add sample refugee metadata
Update dataset structure
Delete file
Delete file
Add sample prwp annotations
Add sample refugee annotations
Add source to subset
Try converted files
Serialize original dict
Normalize metadata json files by wrapping it to a single key-pair
Normalize metadata keys
Add configs to frontmatter
Test upload
initial commit
