smritijha19/Nivi_combined_clarification_set_1
Nivi Clarification Policy Combined v1 Synthetic counterfactual dataset for training a maternal-health clarification policy model. Task Given a maternal-health user query and available context, choose exactly one action: ANSWER_DIRECT ASK_CLARIFICATION ESCALATE_DIRECTLY The model output is a short reasoning trace in <think>...</think> followed by a JSON decision. Language User queries may be in English, Hindi, Assamese, or code-mixed language. The… See the full description on the dataset page: https://huggingface.co/datasets/smritijha19/Nivi_combined_clarification_set_1.
Nivi Clarification Policy Combined v1
Synthetic counterfactual dataset for training a maternal-health clarification policy model.
Task
Given a maternal-health user query and available context, choose exactly one action:
ANSWER_DIRECTASK_CLARIFICATIONESCALATE_DIRECTLY
The model output is a short reasoning trace in <think>...</think> followed by a JSON decision.
Language
User queries may be in English, Hindi, Assamese, or code-mixed language. The reasoning trace and JSON output are in English.
Included families
This combined v1 dataset includes:
- bleeding in pregnancy
- preeclampsia-like headache/swelling symptoms
- abdominal pain in pregnancy
- postpartum discharge/infection concern
- hard negatives / non-triage informational questions
Important caveat
This is a synthetic dataset generated from provisional clarification rubrics and scenario cards. It is not final expert-labeled clinical data. Real Nivi-style examples were used only as style and topic anchors, not as public gold-labeled raw data.
Intended use
Initial SFT baseline for a small LLM clarification policy model.
