Sidharth1743/indicphi
IndicPHI Synthetic Indian clinical documents with PHI/PII span labels for 23 languages (22 scheduled Indian languages + English). All identifiers are synthetic surrogates, not real patients. Two Hub configs at the repo root: Config Files Rows What it is default train.jsonl, eval.jsonl 22,554 / 3,982 Full documents: text, character spans, metadata gliner gliner_train.json, gliner_eval.json 22,889 / 4,054 GLiNER windows: tokenized_text + token ner Also on the… See the full description on the dataset page: https://huggingface.co/datasets/Sidharth1743/indicphi.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face