CoolFace
Datasetpublic

sabin1234/malaria_trend_immunization_measles_rubella_diarrhoea_incidence_fp_modern_health_services_trend

malaria_trend_immunization_measles_rubella_diarrhoea_incidence_fp_modern_health_services_trend_adolescent_abortion_adolescent_anc_phc_orc_patient_types_nurses_reg A merged, ShareGPT-formatted Nepali (नेपाली) health-statistics instruction dataset, combining 11 individual government health data sources into one JSONL file (332 rows). Short name used in this repository: merged_new_health_332_sharegpt_ne Full combined name (all 11 source datasets joined): see title above.… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/malaria_trend_immunization_measles_rubella_diarrhoea_incidence_fp_modern_health_services_trend.

sourceHugging Faceapache-2.0updated 10d agoView on Hugging Face
0likes42downloads
Dataset Card

malariatrendimmunizationmeaslesrubelladiarrhoeaincidencefpmodernhealthservicestrendadolescentabortionadolescentancphcorcpatienttypesnurses_reg

A merged, ShareGPT-formatted Nepali (नेपाली) health-statistics instruction dataset, combining 11 individual government health data sources into one JSONL file (332 rows).

Short name used in this repository: merged_new_health_332_sharegpt_ne Full combined name (all 11 source datasets joined): see title above.

Table of Contents

  1. 1.Overview
  2. 2.Dataset at a Glance
  3. 3.Source Datasets That Were Merged
  4. 4.File Format
  5. 5.Top-Level Schema (Fields)
  6. 6.`conversations` Structure (ShareGPT Format)
  7. 7.`metadata_json` Sub-Fields
  8. 8.Domain / Category / Sub-Domain Breakdown
  9. 9.Behavior Diversity
  10. 10.Question Pattern Diversity
  11. 11.Sample Question–Answer Pairs (One Per Source)
  12. 12.License Breakdown
  13. 13.Text Length Statistics
  14. 14.Coverage: Provinces & Fiscal Years
  15. 15.How to Load the Dataset (Code Demos)
  16. 16.Suggested Use Cases
  17. 17.Known Limitations
  18. 18.Credits / Attribution
  19. 19.Comparison With the Companion Dataset (1,752-Row Merge)

1. Overview

This dataset is a merged collection of 11 separate Nepali-language health-statistics datasets, unified into a single .jsonl file using the ShareGPT conversational format ({"conversations": [{"from": "human", ...}, {"from": "gpt", ...}]}).

Every row is a single-turn, fact-based Question & Answer pair written entirely in Nepali (Devanagari script), sourced from real Nepal government health publications published by the Department of Health Services (DoHS). Each answer restates the exact factual figure, percentage, or rate asked about, making the dataset suitable for training or evaluating models on grounded, factual, Nepali-language question answering in the public-health domain — with a strong focus on maternal, adolescent, and reproductive health, immunization, and disease trend indicators.

The file that was uploaded and analyzed for this README is:

merged_new_health_332_sharegpt_ne.jsonl

2. Dataset at a Glance

PropertyValue
File name analyzedmerged_new_health_332_sharegpt_ne.jsonl
File formatJSON Lines (.jsonl) — one JSON object per line
Total records (rows)332
File size534,569 bytes (≈ 522 KB)
Conversation formatShareGPT (human → gpt, 2 turns per record)
Turns per conversationAlways exactly 2 (1 human question + 1 gpt answer)
LanguageNepali (ne / ISO 639-3 npi)
ScriptDevanagari (Deva)
Task typeinstruction-following
Generation typereal (derived from real published statistics, not synthetic)
Conditionreal
English content allowedFalse — dataset is 100% Nepali
Number of distinct source datasets merged11
Number of distinct source_repo publishers1 (Department of Health Services)
Number of distinct id prefixes (topic buckets)11 (same as source_name — 1:1 in this file)
Duplicate id values0 (all 332 IDs are unique)
Domain (generation_domain)स्वास्थ्य (Health) — 100% of rows
generation_categoryस्वास्थ्य सेवा तथ्यांक (Health service statistics) — 100% of rows (single category)

3. Source Datasets That Were Merged

The dataset name in the title of this README is the combination of the `source_name` value of every dataset merged into this file. The table below shows each individual source, how many rows it contributed, and its license.

#`source_name` (unique ID)Row Count% of TotalLicense
1malaria_trend6820.48%cc-by
2immunization4914.76%cc-by
3measles_rubella3410.24%cc-by
4diarrhoea_incidence3410.24%cc-by
5fp_modern339.94%cc-by
6health_services_trend298.73%cc-by
7adolescent_abortion267.83%cc-by-sa
8adolescent_anc257.53%cc-by-sa
9phc_orc185.42%cc-by
10patient_types103.01%cc-by-sa
11nurses_reg61.81%cc-by
Total332100%

All 11 sources share the same publisher (source_repo): Department of Health Services.

What each source actually covers:

`source_name`Real-world topic covered
malaria_trendMalaria trend indicators over multiple fiscal years, including total national population figures used as denominators
immunizationImmunization programme targets vs. achievements (e.g., BCG and other vaccine coverage)
measles_rubellaMeasles-Rubella surveillance: Non-Measles Non-Rubella (NMNR) incidence counts and rates, by province
diarrhoea_incidenceEstimated population under 5 years at risk of diarrhoea, by province and fiscal year
fp_modernModern family planning method usage (interval methods), by province and fiscal year
health_services_trendMulti-year trend of a health-service indicator (e.g., pill distribution cycles)
adolescent_abortionAdolescent abortion statistics — medical vs. surgical abortion rates/percentages, national level
adolescent_ancAdolescent first Antenatal Care (ANC) visit percentages, including protocol-timed vs. any-time visits
phc_orcPrimary Health Care Outreach Clinic (PHC-ORC) counts and client/service-recipient counts, by province
patient_typesPatient counts by clinical department/type (e.g., general medicine, dermatology)
nurses_regRegistered nurse and AHW (Auxiliary Health Worker) registration counts as of a specific date

4. File Format

  • —Format: JSON Lines (.jsonl) — each line is one independent, valid JSON object (no surrounding array, no trailing commas between lines).
  • —Encoding: UTF-8 (required, since almost all text is Devanagari script).
  • —Conversation schema: ShareGPT-style — a list of turns under the conversations key, alternating human and gpt roles.
  • —Line count = record count: 332 lines = 332 records.

5. Top-Level Schema (Fields)

Every JSON object (row) in the file has exactly the following top-level fields:

FieldTypeAlways PresentDescriptionExample
idstringYesUnique record identifier, formatted as <topic-prefix>_<zero-padded-number>"patient_types_0001"
conversationsarray of objectsYesThe ShareGPT conversation turns (see Section 6)see below
sourcestringYesHuman-readable description of the original data source/publication (constant: "Department of Health Services")"Department of Health Services"
source_namestringYesMachine-readable short key of the source dataset (11 unique values)"malaria_trend"
source_repostringYesGovernment body that published the underlying data (constant)"Department of Health Services"
source_configstringYesDataset configuration name (constant)"default"
source_splitstringYesDataset split label (constant)"train"
source_revisionstringYesVersion/revision tag of the source (constant)"v1"
source_row_idstringYesRow ID within the original source (mirrors id)"patient_types_0001"
languagestringYesISO 639-1 language code (constant)"ne"
language_codestringYesISO 639-3 language code (constant)"npi"
scriptstringYesWriting script (constant)"Deva"
licensestringYesLicense of the underlying source data (2 unique values)"cc-by" or "cc-by-sa"
license_tierstringYesLicense permissiveness tier (constant)"permissive"
task_typestringYesNLP task category (constant)"instruction-following"
generation_typestringYesWhether data is real or synthetically generated (constant)"real"
conditionstringYesData condition flag (constant)"real"
urlstringYesSource URL, if any (empty for every row in this file)""
metadata_jsonstring (JSON-encoded)YesA stringified JSON object with fine-grained metadata — must be parsed separately (see Section 7)"{\"generation_domain\": ...}"
⚠️ Important: metadata_json is stored as an escaped string, not a nested JSON object — you must run json.loads() on it a second time after parsing the outer row.

6. conversations Structure (ShareGPT Format)

Each conversations value is a list of exactly 2 turns:

json
"conversations": [
  {
    "from": "human",
    "value": "आर्थिक वर्ष २०७३/७४ मा सामान्य चिकित्सा बिरामी संख्या कति थियो?"
  },
  {
    "from": "gpt",
    "value": "आर्थिक वर्ष २०७३/७४ मा सामान्य चिकित्सा बिरामी संख्या ४१२०० थियो।"
  }
]
Turn FieldTypeDescription
fromstringSpeaker role — always either "human" or "gpt"
valuestringThe message text in Nepali (Devanagari script)
PropertyValue
Turns per conversation2 (always)
Multi-turn conversations present?No — 100% of records are single-turn Q&A
Roles presenthuman (332 messages), gpt (332 messages)

7. metadata_json Sub-Fields

Once parsed (json.loads(row["metadata_json"])), each metadata object contains the following keys:

Sub-fieldTypeUnique ValuesDescriptionExample
generation_domainstring1Top-level subject domain"स्वास्थ्य" (Health)
generation_categorystring1Mid-level topic category (constant in this file)"स्वास्थ्य सेवा तथ्यांक" (Health service statistics)
generation_sub_domainstring2Fine-grained sub-topic grouping (see Section 8)"स्वास्थ्य सूचक"
behaviorstring1The intended model behavior label"तथ्यमा आधारित उत्तर दिने" (give a fact-based answer)
behavior_definitionstring1Full description of the required behavior"मूल तथ्य कायम राख्दै शुद्ध नेपालीमा उत्तर दिने।" (preserve the core fact and answer in pure Nepali)
question_typestring1Category of question"तथ्यमा आधारित प्रश्न" (fact-based question)
content_languagestring1Content language name"नेपाली"
content_scriptstring1Content script name"देवनागरी"
english_content_allowedboolean1Whether English text is permitted in answersfalse
Note: in this 332-row file, only `generation_sub_domain` varies (2 values) — every other metadata_json field, including generation_category, is constant. This is a narrower, more topically-focused merge than the companion 1,752-row dataset.

8. Domain / Category / Sub-Domain Breakdown

8.1 generation_category (1 unique value — constant)

`generation_category` (Nepali)English MeaningRow Count% of Total
स्वास्थ्य सेवा तथ्यांकHealth service statistics332100%

8.2 generation_sub_domain (2 unique values)

`generation_sub_domain` (Nepali)English MeaningRow Count% of Total
स्वास्थ्य सूचकHealth indicators15847.59%
जनसंख्या तथा सेवाPopulation & services17452.41%

8.3 source_name ↔ generation_sub_domain Cross-Reference

`source_name`Mapped `generation_sub_domain`
patient_typesस्वास्थ्य सूचक (Health indicators)
malaria_trendस्वास्थ्य सूचक (Health indicators)
adolescent_abortionस्वास्थ्य सूचक (Health indicators)
adolescent_ancस्वास्थ्य सूचक (Health indicators)
health_services_trendस्वास्थ्य सूचक (Health indicators)
measles_rubellaजनसंख्या तथा सेवा (Population & services)
nurses_regजनसंख्या तथा सेवा (Population & services)
phc_orcजनसंख्या तथा सेवा (Population & services)
diarrhoea_incidenceजनसंख्या तथा सेवा (Population & services)
immunizationजनसंख्या तथा सेवा (Population & services)
fp_modernजनसंख्या तथा सेवा (Population & services)

8.4 id Prefix Breakdown

Unlike the companion 1,752-row dataset, this file's id prefixes map 1:1 onto source_name (no further topic-splitting occurs):

`id` prefix = `source_name`Row Count
malaria_trend68
immunization49
measles_rubella34
diarrhoea_incidence34
fp_modern33
health_services_trend29
adolescent_abortion26
adolescent_anc25
phc_orc18
patient_types10
nurses_reg6

9. Behavior Diversity

Like its companion dataset, this file uses a structured metadata schema that formally labels every row with a behavior and behavior_definition.

9.1 Formal behavior Label (from metadata_json)

`behavior` (Nepali)English MeaningRow Count% of Total
तथ्यमा आधारित उत्तर दिनेGive a fact-based answer332100%

`behavior_definition` (constant across all rows):

NepaliEnglish Translation
मूल तथ्य कायम राख्दै शुद्ध नेपालीमा उत्तर दिने।Answer in pure Nepali while preserving the core underlying fact.

➡️ Exactly as in the companion dataset, the formal metadata is intentionally uniform: 100% of rows are tagged with a single behavior class — "answer factually, in Nepali, without altering the underlying statistic." This is a single-behavior, high-precision factual QA dataset.

9.2 Practical Answer-Behavior Patterns Observed

Even though the metadata behavior tag never changes, the gpt answers exhibit several distinct, recurring answering behaviors depending on the underlying indicator type:

Observed Answer BehaviorDescriptionWhich sources show it most
Direct value restatementAnswer restates the question as a full sentence with the number substituted inAll 11 sources
Percentage-with-raw-count reportingAnswer reports a percentage and the underlying raw count together, e.g. 11.6% (6196)adolescent_abortion
Pure percentage/rate reportingAnswer reports only a computed percentage or rate, no raw countadolescent_anc, measles_rubella
Target-vs-achievement pairingQuestion/answer pairs exist for both a "target" and an "achievement" value for the same indicatorimmunization
Multi-year trend-point reportingAnswer reports a single year's value drawn from a repeating multi-year indicator seriesmalaria_trend, health_services_trend
Date-anchored (not fiscal-year) reportingAnswer is anchored to a specific Bikram Sambat calendar date rather than a fiscal yearnurses_reg

10. Question Pattern Diversity

Every human question in this dataset is unique as an exact string (332 out of 332 questions are exact-text-unique). To measure structural pattern diversity (i.e., how many distinct question templates exist once fiscal years, provinces, and numbers are treated as variables), all Devanagari/Arabic digits were stripped and the remaining unique phrasing was counted per source.

10.1 Pattern Diversity Per Source

`source_name`Total QuestionsExact-Text-UniqueNormalized Unique PatternsPattern Diversity (Normalized / Total)
patient_types101010100.0%
immunization49493163.3%
malaria_trend68682435.3%
health_services_trend29291137.9%
measles_rubella34341029.4%
diarrhoea_incidence34341029.4%
fp_modern3333927.3%
adolescent_abortion2626830.8%
adolescent_anc2525728.0%
nurses_reg666100.0%
phc_orc1818633.3%
Overall dataset33233213239.8%
Interpretation: patient_types and nurses_reg ask a genuinely different question every single time (every row is a distinct patient/department type or a distinct registration category — 100% template diversity). immunization and malaria_trend reuse a moderate number of templates (24–31) across different vaccines/years. Sources like adolescent_anc, fp_modern, and phc_orc reuse a small number of question templates (6–9) and vary mainly the province or fiscal year.

10.2 Example Question Templates — immunization (highest-diversity multi-row source)

#Question Template (numbers stripped)
1Fiscal year # — what was the target for [vaccine]?
2Fiscal year # — what was the achievement for [vaccine]?
3Fiscal year # — what was the coverage percentage for [vaccine]?
...(pattern repeats per vaccine: BCG, DPT-HepB-Hib, OPV, Measles-Rubella, etc.)

10.3 Example Question Templates — Low-Diversity Sources

SourceDistinct TemplatesExample Template
adolescent_anc7"Fiscal year # — nationally, what was the percentage of adolescents' first ANC visit (protocol-timed)?"
fp_modern9"Fiscal year # — Province # — how many modern family-planning interval-method users were there?"
phc_orc6"Fiscal year # — Province # — how many Primary Health Care Outreach Clinics were there?"

11. Sample Question–Answer Pairs (One Per Source)

`source_name`Sample Question (Nepali)Sample Answer (Nepali)
patient_typesआर्थिक वर्ष २०७३/७४ मा सामान्य चिकित्सा बिरामी संख्या कति थियो?आर्थिक वर्ष २०७३/७४ मा सामान्य चिकित्सा बिरामी संख्या ४१२०० थियो।
malaria_trendआर्थिक वर्ष २०७१/७२ मा कुल जनसंख्या कति थियो?आर्थिक वर्ष २०७१/७२ मा कुल जनसंख्या १३४५५००० थियो।
adolescent_abortionआर्थिक वर्ष २०७३/७४ मा राष्ट्रिय मा औषधिबाट गर्भपतन कति थियो?आर्थिक वर्ष २०७३/७४ मा राष्ट्रिय मा औषधिबाट गर्भपतन ११.६% (६१९६) थियो।
adolescent_ancआर्थिक वर्ष २०७३/७४ मा राष्ट्रिय मा किशोरीको पहिलो गर्भवती जाँच (जुनसुकै समय) प्रतिशत कति थियो?आर्थिक वर्ष २०७३/७४ मा राष्ट्रिय मा किशोरीको पहिलो गर्भवती जाँच (जुनसुकै समय) प्रतिशत १९.२ थियो।
health_services_trendआर्थिक वर्ष २०७१/७२ मा चक्की वितरण (चक्र संख्या) कति थियो?आर्थिक वर्ष २०७१/७२ मा चक्की वितरण (चक्र संख्या) ८६६८८१ थियो।
measles_rubellaप्रदेश १ मा एनएमएनआर घटना कति थिए?प्रदेश १ मा एनएमएनआर घटना १७० थिए।
nurses_reg२०७४ कात्तिक २६ सम्म नर्स दर्ता संख्या कति थियो?२०७४ कात्तिक २६ सम्म नर्स दर्ता संख्या ४४७३७ थियो।
phc_orcआर्थिक वर्ष २०७३/७४ मा प्रदेश १ मा प्राथमिक स्वास्थ्य हेरचाह आउटरीच क्लिनिक संख्या कति थियो?आर्थिक वर्ष २०७३/७४ मा प्रदेश १ मा प्राथमिक स्वास्थ्य हेरचाह आउटरीच क्लिनिक संख्या २४७९२ थियो।
diarrhoea_incidenceआर्थिक वर्ष २०७३/७४ मा प्रदेश १ मा झाडापखालाको जोखिममा रहेका ५ वर्षमुनिका अनुमानित जनसंख्या कति थियो?आर्थिक वर्ष २०७३/७४ मा प्रदेश १ मा झाडापखालाको जोखिममा रहेका ५ वर्षमुनिका अनुमानित जनसंख्या ४९४३०१ थियो।
immunizationआर्थिक वर्ष २०७३/७४ मा बीसीजी को लक्ष्य कति थियो?आर्थिक वर्ष २०७३/७४ मा बीसीजी को लक्ष्य ६२३९२९ थियो।
fp_modernआर्थिक वर्ष २०७३/७४ मा प्रदेश १ मा परिवार नियोजन अन्तराल विधि प्रयोगकर्ता कति थिए?आर्थिक वर्ष २०७३/७४ मा प्रदेश १ मा परिवार नियोजन अन्तराल विधि प्रयोगकर्ता २६६०४२ थिए।

12. License Breakdown

`license`Row Count% of TotalAffected `source_name` values
cc-by27181.63%malaria_trend, immunization, measles_rubella, diarrhoea_incidence, fp_modern, health_services_trend, phc_orc, nurses_reg
cc-by-sa6118.37%adolescent_abortion, adolescent_anc, patient_types
`license_tier`Row Count
permissive332 (100%)
⚠️ cc-by-sa (ShareAlike) is more restrictive than plain cc-by: any derivative dataset built substantially from adolescent_abortion, adolescent_anc, or patient_types rows should be released under a compatible ShareAlike license.

13. Text Length Statistics

MetricHuman QuestionGPT Answer
Minimum length (characters)2715
Maximum length (characters)103106
Mean length (characters)65.466.5

14. Coverage: Provinces & Fiscal Years

Coverage DimensionValue
Provinces referencedप्रदेश १ – प्रदेश ७ (all 7 provinces of Nepal)
Distinct Nepali fiscal years (B.S.) referenced3 — २०७१/७२, २०७२/७३, २०७३/७४
Non-fiscal-year date referencenurses_reg uses a specific B.S. calendar date (२०७४ कात्तिक २६) instead of a fiscal year

Fiscal-year mention frequency

Fiscal Year (B.S.)Approx. Gregorian EquivalentMentions
२०७१/७२2014/1534
२०७२/७३2015/1632
२०७३/७४2016/17213
This dataset's fiscal-year coverage is narrower and more recent than the companion 1,752-row dataset (which spans 9 fiscal years back to २०६५/६६) — over 64% of this file's fiscal-year-tagged questions concern २०७३/७४ (2016/17) alone.

15. How to Load the Dataset (Code Demos)

15.1 Plain Python (no dependencies)

python
import json

file_path = "merged_new_health_332_sharegpt_ne.jsonl"

records = []
with open(file_path, "r", encoding="utf-8") as f:
    for line in f:
        line = line.strip()
        if not line:
            continue
        records.append(json.loads(line))

print(f"Total records loaded: {len(records)}")

# Inspect the first record
first = records[0]
print("ID:", first["id"])
print("Source:", first["source_name"])

question = first["conversations"][0]["value"]
answer = first["conversations"][1]["value"]
print("Q:", question)
print("A:", answer)

# metadata_json is a STRING — parse it separately
meta = json.loads(first["metadata_json"])
print("Sub-domain:", meta["generation_sub_domain"])

15.2 Using pandas

python
import json
import pandas as pd

file_path = "merged_new_health_332_sharegpt_ne.jsonl"

# Read line-by-line (jsonl -> DataFrame)
rows = []
with open(file_path, "r", encoding="utf-8") as f:
    for line in f:
        line = line.strip()
        if line:
            rows.append(json.loads(line))

df = pd.DataFrame(rows)

# Expand the stringified metadata_json into real columns
meta_df = pd.json_normalize(df["metadata_json"].apply(json.loads))
df = pd.concat([df.drop(columns=["metadata_json"]), meta_df], axis=1)

# Pull question / answer text out of the conversations column
df["question"] = df["conversations"].apply(lambda c: c[0]["value"])
df["answer"]   = df["conversations"].apply(lambda c: c[1]["value"])

print(df.shape)
print(df[["id", "source_name", "generation_sub_domain", "question", "answer"]].head())

# Example: filter to a single source
malaria = df[df["source_name"] == "malaria_trend"]
print(f"{len(malaria)} malaria-trend rows")

15.3 Using pandas' built-in JSON-lines reader (shortcut)

python
import pandas as pd

df = pd.read_json("merged_new_health_332_sharegpt_ne.jsonl", lines=True)
print(df.shape)
print(df.columns.tolist())

15.4 Using the Hugging Face datasets library

python
from datasets import load_dataset

dataset = load_dataset(
    "json",
    data_files="merged_new_health_332_sharegpt_ne.jsonl",
    split="train",
)

print(dataset)
print(dataset[0])

# Filter by source_name
immunization_only = dataset.filter(lambda row: row["source_name"] == "immunization")
print(len(immunization_only))

15.5 Streaming line-by-line (for very large files / low memory)

python
import json

def iter_records(path):
    with open(path, "r", encoding="utf-8") as f:
        for line in f:
            line = line.strip()
            if line:
                yield json.loads(line)

for i, rec in enumerate(iter_records("merged_new_health_332_sharegpt_ne.jsonl")):
    q = rec["conversations"][0]["value"]
    a = rec["conversations"][1]["value"]
    # process record here
    if i < 3:
        print(f"[{rec['id']}] Q: {q}\n[{rec['id']}] A: {a}\n")

15.6 Converting to OpenAI-style / Alpaca-style fine-tuning format

python
import json

def to_openai_chat_format(record):
    q = record["conversations"][0]["value"]
    a = record["conversations"][1]["value"]
    return {
        "messages": [
            {"role": "user", "content": q},
            {"role": "assistant", "content": a},
        ]
    }

with open("merged_new_health_332_sharegpt_ne.jsonl", "r", encoding="utf-8") as f_in, \
     open("openai_format_332.jsonl", "w", encoding="utf-8") as f_out:
    for line in f_in:
        rec = json.loads(line)
        f_out.write(json.dumps(to_openai_chat_format(rec), ensure_ascii=False) + "\n")

15.7 Merging this file with the companion 1,752-row dataset

python
import json

files = [
    "merged_health_dataset_sharegpt_ne__1_.jsonl",
    "merged_new_health_332_sharegpt_ne.jsonl",
]

combined = []
seen_ids = set()
for path in files:
    with open(path, "r", encoding="utf-8") as f:
        for line in f:
            line = line.strip()
            if not line:
                continue
            rec = json.loads(line)
            if rec["id"] in seen_ids:
                raise ValueError(f"Duplicate id across files: {rec['id']}")
            seen_ids.add(rec["id"])
            combined.append(rec)

print(f"Combined total: {len(combined)} rows")

with open("combined_health_dataset_2084.jsonl", "w", encoding="utf-8") as f_out:
    for rec in combined:
        f_out.write(json.dumps(rec, ensure_ascii=False) + "\n")

16. Suggested Use Cases

Use CaseWhy This Dataset Fits
Fine-tuning a Nepali-language factual QA model on maternal/adolescent/reproductive health51 rows (adolescent_abortion + adolescent_anc) focus specifically on adolescent reproductive-health indicators
Evaluating Nepali LLM factual recall on immunization and vaccine-programme statisticsimmunization alone provides 49 target-vs-achievement Q&A pairs across multiple vaccines
Building a Nepali retrieval-augmented-generation (RAG) benchmark on recent (FY 2073/74) health dataOver 64% of fiscal-year-tagged rows concern the single most recent fiscal year in this file
Studying percentage-with-raw-count answer formatting in Nepali (e.g., 11.6% (6196))adolescent_abortion rows explicitly combine a percentage and its underlying count in one answer
Extending/merging with the companion 1,752-row dataset for a larger unified Nepali health-QA corpusBoth files share the identical schema, field set, and id uniqueness — see Section 15.7

17. Known Limitations

  • —Single-turn only: No multi-turn or follow-up-question conversations are present (100% of rows have exactly 2 turns).
  • —Uniform behavior/category labels: The behavior, question_type, and generation_category metadata fields carry no diversity in this file (always the same value) — genuine diversity exists only in generation_sub_domain (2 values) and in the underlying question phrasing/topic.
  • —Template reuse in several sources: Sources such as adolescent_anc, fp_modern, and phc_orc reuse a small number of question templates (6–9) and vary mainly the province or fiscal year being asked about — see Section 10.
  • —Small sources: nurses_reg (6 rows) and patient_types (10 rows) are very small compared to malaria_trend (68 rows), which may cause class imbalance if this dataset is used for per-source stratified evaluation.
  • —Mixed licensing: 18.4% of rows (adolescent_abortion, adolescent_anc, patient_types) are licensed cc-by-sa (ShareAlike) rather than plain cc-by.
  • —No English content: The dataset is 100% Nepali; it is not suitable as-is for multilingual or English-Nepali parallel-text tasks.
  • —Narrow fiscal-year range: Only 3 fiscal years are referenced (२०७१/७२–२०७३/७४, roughly 2014–2017 Gregorian), a much narrower time span than the companion dataset's 9-year range.

18. Credits / Attribution

  • —Primary publisher: Department of Health Services (DoHS), Nepal — 100% of records in this file.
  • —Format: Compiled and merged into ShareGPT conversational JSONL format for instruction-tuning / QA use.

If you use this dataset, please retain attribution to the original Department of Health Services publications listed in the source field of each record.


19. Comparison With the Companion Dataset (1,752-Row Merge)

This dataset is designed to be schema-compatible with the earlier, larger merge (merged_health_dataset_sharegpt_ne__1_.jsonl). Key differences:

PropertyThis dataset (332 rows)Companion dataset (1,752 rows)
Number of source datasets merged1113
Publishers (source_repo)1 (DoHS only)2 (DoHS + MoHP)
generation_category values1 (constant)9
generation_sub_domain values27
id prefix granularitySame as source_name (11 prefixes)Finer than source_name (20 prefixes)
License types presentcc-by, cc-by-sacc-by, License not specified
Fiscal years covered3 (२०७१/७२–२०७३/७४)9 (२०६५/६६–२०७३/७४)
Overall normalized question-pattern diversity39.8%77.3%
Dominant topic focusImmunization, malaria, adolescent/maternal health, family planningCommunicable disease outbreaks, FCHV workforce, hospital resources

Both files use the identical top-level schema and `metadata_json` sub-schema, so they can be safely concatenated (see the merge code snippet in Section 15.7) as long as id uniqueness across both files is verified first.