sabin1234/drug_poisoning_nepali_sharegpt
README — drug_poisoning_nepali_sharegpt_cleaned.jsonl This README file provides detailed information about the drug_poisoning_nepali_sharegpt_cleaned.jsonl dataset — file format, schema, source, subject matter, question pattern diversity, answer behaviour diversity, geographic/temporal coverage, and statistical analysis, all presented in tables. 1. General File Information Detail Value File name drug_poisoning_nepali_sharegpt_cleaned.jsonl Format… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/drug_poisoning_nepali_sharegpt.
README — drug_poisoning_nepali_sharegpt_cleaned.jsonl
This README file provides detailed information about the drug_poisoning_nepali_sharegpt_cleaned.jsonl dataset — file format, schema, source, subject matter, question pattern diversity, answer behaviour diversity, geographic/temporal coverage, and statistical analysis, all presented in tables.
1. General File Information
2. Dataset Domain Overview
3. JSON Schema (fields in each record)
4. Value Diversity of Metadata Fields (Constant vs Variable Fields)
5. Basic Statistics
6. Question Pattern Diversity
All questions in the dataset are surface-level distinct (56,448 out of 56,448 are unique), but their underlying structure (template) is based on a limited number of paraphrase variants. Even after normalizing (removing numbers like year, population, ID), many different sentence structures are found, whose main characteristics are shown in the tables below.
6.1 Question Ending Patterns — 19 distinct patterns total
6.2 Diversity by Question Sentence Type (Sentence Mood)
6.3 Diversity by Question Opening Structure
6.4 Information Pattern within Questions
Note: Every question in the dataset follows the same "single-fact lookup" pattern — asking for the population and mortality rate category of a specific state + district + year. There are no complex questions such as comparisons, trends, or multi-year queries.
6.5 Synonym / Lexical Variation
Different Nepali words used to express the same concept (paraphrase diversity):
7. Answer Behaviour Diversity
7.1 Answer Opening Patterns (Top 20 — out of 3,330 distinct openings total)
7.2 Answer Content Structure
7.3 Distribution of Mortality Rate Categories
These categories resemble CDC's standard "age-adjusted death rate range/category" format — where lower-numbered ranges (such as<2or2-3.9) are more common and higher ranges (up to28-29.9) appear only in rare counties.
8. Geographic Coverage
8.1 Top 20 States by Presence
9. Temporal Coverage
10. Sample Records (showing diverse styles)
11. Data Quality Summary
12. License & Usage
13. Summary
This dataset is built on county-level population and age-adjusted drug/substance poisoning mortality rate statistics sourced from the US CDC, and consists of 56,448 synthetic single-turn question-answer pairs prepared in Nepali. It includes multi-dimensional diversity in lexical choice, syntactic structure, and sentence mood (interrogative vs. imperative), but the subject matter is always limited to the same kind of "one state + one district + one year" population and mortality-rate-category lookup pattern (there are no comparative or multi-year questions).
