mukuls9971/address-benchmark-v1
Indian Address Benchmark Dataset v1 Mixed benchmark dataset for Indian-address tagging built from synthetic data plus public upstream datasets. Repository Dataset repo: mukuls9971/address-benchmark-v1 Train split: 26728 Validation split: 6158 Test split: 1410 Files train.jsonl validation.jsonl test.jsonl report.json Notes Generated and published by the pii-model-oss workflow. Upstream datasets used to assemble benchmark variants… See the full description on the dataset page: https://huggingface.co/datasets/mukuls9971/address-benchmark-v1.
031
Indian Address Benchmark Dataset v1
Mixed benchmark dataset for Indian-address tagging built from synthetic data plus public upstream datasets.
Repository
- Dataset repo:
mukuls9971/address-benchmark-v1 - Train split:
26728 - Validation split:
6158 - Test split:
1410
Files
train.jsonlvalidation.jsonltest.jsonlreport.json
Notes
- Generated and published by the
pii-model-ossworkflow. - Upstream datasets used to assemble benchmark variants retain their own licenses.
Warnings
- LinCE train/dev could not be fetched from the original host; used CodeMixBench ner_hineng test as a held-out-only fallback.
