CoolFace
Datasetpublic

sabin1234/drug_poisoning_nepali_sharegpt

README — drug_poisoning_nepali_sharegpt_cleaned.jsonl This README file provides detailed information about the drug_poisoning_nepali_sharegpt_cleaned.jsonl dataset — file format, schema, source, subject matter, question pattern diversity, answer behaviour diversity, geographic/temporal coverage, and statistical analysis, all presented in tables. 1. General File Information Detail Value File name drug_poisoning_nepali_sharegpt_cleaned.jsonl Format… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/drug_poisoning_nepali_sharegpt.

sourceHugging Faceapache-2.0updated 19d agoView on Hugging Face
0likes47downloads
Dataset Card

README — drug_poisoning_nepali_sharegpt_cleaned.jsonl

This README file provides detailed information about the drug_poisoning_nepali_sharegpt_cleaned.jsonl dataset — file format, schema, source, subject matter, question pattern diversity, answer behaviour diversity, geographic/temporal coverage, and statistical analysis, all presented in tables.


1. General File Information

DetailValue
File namedrug_poisoning_nepali_sharegpt_cleaned.jsonl
FormatJSON Lines (.jsonl) — each line is an independent JSON object
Total records (rows)56,448
File size~77 MB
Conversation structureShareGPT-style (human and gpt roles inside a conversations array)
Number of turns per record2 (1 human + 1 gpt) — 100% single-turn
LanguageNepali (ne / npi)
ScriptDevanagari (Deva)
Duplicate records0 — all ids, questions, and answers are completely unique

2. Dataset Domain Overview

AspectDetail
Original source dataCenters for Disease Control and Prevention (CDC), USA
DomainCounty-level population and "Drug Poisoning Mortality" (age-adjusted death rate)
Variables measured(1) Annual population of the relevant county/district, (2) Age-adjusted mortality rate category (range)
Geographic unitUS State + County/District
Time range1999 to 2016 (18 years)
Data typeSynthetic — question-answer pairs generated in Nepali by a model, based on real CDC statistics
Generating modelhimalaya-gemma-4-e2b-it
Task typeInstruction-following (single-fact retrieval / lookup QA)

3. JSON Schema (fields in each record)

Field nameData typeDescriptionExample value
idstringUnique record identifierdrug_poisoning_sft_000001
conversationsarrayArray containing 2 objects with human and gpt rolessee below
conversations[0].fromstringAlways humanhuman
conversations[0].valuestringUser's question (in Nepali)—
conversations[1].fromstringAlways gptgpt
conversations[1].valuestringModel's answer (in Nepali)—
sourcestringType of data originsynthetic
source_namestringOriginal data-providing organizationCenters for Disease Control and Prevention
source_repostringSource repository namelocal_generation
source_configstringDataset configuration/domain namedrug_poisoning_mortality
source_splitstringData splittrain
source_revisionstringVersion numberv1.0
source_row_idstringRow sequence number of the original table1, 2, 3 ...
languagestringLanguage code (short)ne
language_codestringLanguage code (ISO)npi
scriptstringScriptDeva
licensestringLicense detailPublic Domain - U.S. Government
license_tierstringLicense classpermissive
task_typestringTask typeinstruction-following
generation_typestringGeneration methodsynthetic
conditionstringData conditionmodel-generated
urlstringSource URL (empty here)""
metadata_jsonstring (JSON-encoded)Additional metadata (domain, model, generation_number){"domain":"drug_poisoning_mortality","model":"gemini-3.5-flash-lite","generation_number":1}

4. Value Diversity of Metadata Fields (Constant vs Variable Fields)

FieldNumber of unique valuesStatus
id56,448 (all different)Fully variable
source1 (synthetic)Constant
source_name1 (Centers for Disease Control and Prevention)Constant
source_repo1 (local_generation)Constant
source_config1 (drug_poisoning_mortality)Constant
source_split1 (train)Constant
source_revision1 (v1.0)Constant
source_row_id56,448 (all different, sequential)Variable
language1 (ne)Constant
language_code1 (npi)Constant
script1 (Deva)Constant
license1 (Public Domain - U.S. Government)Constant
license_tier1 (permissive)Constant
task_type1 (instruction-following)Constant
generation_type1 (synthetic)Constant
condition1 (model-generated)Constant
metadata_json → model1 (himalaya-gemma-4-e2b-it)Constant
metadata_json → domain1 (drug_poisoning_mortality)Constant
metadata_json → generation_number56,448 (all different)Variable

5. Basic Statistics

MetricHuman (question)GPT (answer)
Total count56,44856,448
Unique text count56,448 (100%)56,448 (100%)
Average character length132.8140.1
Minimum character length117108
Maximum character length180188
Average word count20.420.8
Minimum word count1717
Maximum word count2930

6. Question Pattern Diversity

All questions in the dataset are surface-level distinct (56,448 out of 56,448 are unique), but their underlying structure (template) is based on a limited number of paraphrase variants. Even after normalizing (removing numbers like year, population, ID), many different sentence structures are found, whose main characteristics are shown in the tables below.

6.1 Question Ending Patterns — 19 distinct patterns total

No.Ending patternCountPercentage
1...what was the category?9,16716.24%
2...which category was it in?4,7638.44%
3...which range was it in?4,7148.35%
4...which category is it in?4,5378.04%
5Identify the range.2,4154.28%
6...which range does it show?2,4024.26%
7...which class did it fall into?2,3854.23%
8...what range was it placed in?2,3754.21%
9...what does the record show?2,3554.17%
10Also state the range.2,3464.16%
11...which range did it fall into?2,3464.16%
12...what was the population?2,3384.14%
13State the rate range.2,3334.13%
14...what was the classification?2,3314.13%
15...which level was it at?2,3054.08%
16...what category was recorded?2,2043.90%
17...which range were they in?2,2023.90%
18...which category is shown?2,1923.88%
19...mentioned in the statistics.7381.31%

6.2 Diversity by Question Sentence Type (Sentence Mood)

Sentence typeCountPercentageCharacteristic
Interrogative (ends in ?)48,61686.11%Questions like "...how much?", "...which category was it in?"
Imperative/Request (ends in ।)7,83213.89%Requests like "...state it", "...identify it"

6.3 Diversity by Question Opening Structure

Opening typeCountPercentageExample
Year/time-first24,08442.67%"In the year 2011...", "For the year 1999..."
State/district-first32,36457.33%"In the state of Texas...", "Autauga district..."

6.4 Information Pattern within Questions

Fact-requesting patternPresent in how many questionsPercentage
Population asked56,448100%
Mortality rate category asked56,448100%
Both facts asked together (compound question)56,448100%
Comparison/trend question00%
Note: Every question in the dataset follows the same "single-fact lookup" pattern — asking for the population and mortality rate category of a specific state + district + year. There are no complex questions such as comparisons, trends, or multi-year queries.

6.5 Synonym / Lexical Variation

Different Nepali words used to express the same concept (paraphrase diversity):

ConceptDifferent words/phrases used
"category/level"श्रेणी, वर्ग, दायरा, तह, वर्गीकरण (category, class, range, level, classification)
"year"वर्ष, साल, समय (year, year, time)
"statistics/record"तथ्याङ्क, अभिलेख, डेटा, विवरण (statistics, record, data, detail)
"asking verb"कति छ, के थियो, बताउनुहोस्, पहिचान गर्नुहोस्, देखाउँछ, उल्लेख छ (how much is it, what was it, state it, identify it, shows, is mentioned)
"was/is" (tense)Both past tense (थियो, परेको थियो) and present tense (छ, देखाउँछ) are used

7. Answer Behaviour Diversity

7.1 Answer Opening Patterns (Top 20 — out of 3,330 distinct openings total)

Opening phraseCountPercentage
Based on the record...2,2714.02%
In that year...2,2614.01%
In that year...2,1863.87%
In 2010...5430.96%
In 1999...5210.92%
In 2009...5190.92%
In 2006...5180.92%
In 2013...5150.91%
In 2002...5130.91%
In 2003...5120.91%
In 2001...5000.89%
In 2005...4990.88%
In 2016...4970.88%
In 2012...4940.87%
In 2015...4900.87%
In 2007...4820.85%
In 2014...4780.85%
In 2008...4750.84%
In 2011...4610.82%
In 2004...4560.81%

7.2 Answer Content Structure

Structural elementDescriptionPresence
Population figureAlways a concrete number (e.g. 42963)100% of answers
Mortality rate categoryAlways a numeric range (e.g. 2-3.9) or a bound like <2100% of answers
Sentences per answerUsually 2 sentences (one for population, one for mortality rate); sometimes both combined in 1 sentenceVariable
Past/present tense usageAnswers match the tense of the question, using both past and presentMixed
Filler/Framing phrasesVarious introducers such as "based on the record", "the statistics show", "according to the details"High diversity

7.3 Distribution of Mortality Rate Categories

Category (rate per 100,000 population)Estimated count (based on observed samples)
< 2Present in high numbers
2 – 3.9High
4 – 5.9Most common overall
6 – 7.9High
8 – 9.9Moderate-high
10 – 11.9Moderate
12 – 13.9Moderate
14 – 15.9Low
16 – 17.9Low
18 – 19.9Rare
20 – 21.9Rare
22 – 23.9Rare
24 – 25.9Rare
26 – 27.9Very rare
28 – 29.9Very rare
These categories resemble CDC's standard "age-adjusted death rate range/category" format — where lower-numbered ranges (such as <2 or 2-3.9) are more common and higher ranges (up to 28-29.9) appear only in rare counties.

8. Geographic Coverage

8.1 Top 20 States by Presence

No.StateMentions in questions (approx.)
1Texas1,559
2North Carolina964
3Georgia961
4Virginia889
5Kentucky746
6South Dakota739
7Missouri733
8Illinois653
9Kansas657
10Iowa611
11West Virginia570
12Tennessee561
13Nebraska554
14Indiana544
15Mississippi533
16Minnesota530
17Ohio524
18North Dakota518
19Michigan517
20Arkansas498
Overall geographic statisticValue
Total distinct states found30+ (abbreviated US state names)
Total distinct county/district names (word-based estimate)1,650+
Geographic levelUS State + County level

9. Temporal Coverage

MetricValue
Years used1999, 2000, 2001, 2002, 2003, 2004, 2005, 2006, 2007, 2008, 2009, 2010, 2011, 2012, 2013, 2014, 2015, 2016
Total distinct years18
Most frequently appearing years (example)2007, 2012, 2015 (each ~3,138 times)
Time range typeYearly snapshot; no trend/multi-year questions

10. Sample Records (showing diverse styles)

IDQuestion styleSample questionSample answer
drugpoisoningsft_000001State + District + Year, interrogative, both facts togetherWhat was the population of Autauga district in the state of Alabama in the year 1999, and which category is the age-adjusted mortality rate from drug poisoning in?The population of Autauga district in 1999 was 42963. The age-adjusted mortality rate category from drug poisoning was 2-3.9.
drugpoisoningsft_015001Year-first + imperativeFor the year 2013, identify the population of Rockcastle district in the state of Kentucky and the range of the age-adjusted mortality rate from drug poisoning.The population of Rockcastle district in 2013 was 16744. The age-adjusted mortality rate category from drug poisoning was 24-25.9.
drugpoisoningsft_025001Comma-structured, two-part questionWhat was the population of Los Alamos district in the state of New Mexico in 2000, and which range was the age-adjusted mortality rate from drug poisoning in?The record shows a population of 18263 for Los Alamos district, New Mexico in 2000, and places the age-adjusted mortality rate from drug poisoning in the 6-7.9 category.
drugpoisoningsft_045001District-first reversed orderWhat population did Covington district cover in 2011, and what was the category of the age-adjusted mortality rate from drug poisoning in that district of Mississippi state?In 2011, the population for Covington district, Mississippi was recorded as 19509, and the age-adjusted mortality rate from drug poisoning is recorded in the 8-9.9 range.

11. Data Quality Summary

Check pointResult
Duplicate id0 (all unique)
Duplicate questions (verbatim)0 (all unique)
Duplicate answers (verbatim)0 (all unique)
Turns per recordAlways 2 (human + gpt)
Empty/null fieldsurl is always an empty string; other fields are filled
Language uniformity100% Nepali (Devanagari)
Metadata uniformityMost fields are constant (since this is a single-domain, single-source dataset)

12. License & Usage

AspectDetail
LicensePublic Domain – U.S. Government
License tierPermissive (free to use)
Suitable usesNepali-language instruction-following / QA model fine-tuning, training single-fact retrieval capability, developing numeric-statistics comprehension ability
LimitationsSince the dataset covers only a single domain (drug poisoning mortality), it needs to be combined with additional data for general knowledge or multi-domain training

13. Summary

This dataset is built on county-level population and age-adjusted drug/substance poisoning mortality rate statistics sourced from the US CDC, and consists of 56,448 synthetic single-turn question-answer pairs prepared in Nepali. It includes multi-dimensional diversity in lexical choice, syntactic structure, and sentence mood (interrogative vs. imperative), but the subject matter is always limited to the same kind of "one state + one district + one year" population and mortality-rate-category lookup pattern (there are no comparative or multi-year questions).