sabin1234/malaria_trend_immunization_measles_rubella_diarrhoea_incidence_fp_modern_health_services_trend
malaria_trend_immunization_measles_rubella_diarrhoea_incidence_fp_modern_health_services_trend_adolescent_abortion_adolescent_anc_phc_orc_patient_types_nurses_reg A merged, ShareGPT-formatted Nepali (नेपाली) health-statistics instruction dataset, combining 11 individual government health data sources into one JSONL file (332 rows). Short name used in this repository: merged_new_health_332_sharegpt_ne Full combined name (all 11 source datasets joined): see title above.… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/malaria_trend_immunization_measles_rubella_diarrhoea_incidence_fp_modern_health_services_trend.
malariatrendimmunizationmeaslesrubelladiarrhoeaincidencefpmodernhealthservicestrendadolescentabortionadolescentancphcorcpatienttypesnurses_reg
A merged, ShareGPT-formatted Nepali (नेपाली) health-statistics instruction dataset, combining 11 individual government health data sources into one JSONL file (332 rows).
Short name used in this repository: merged_new_health_332_sharegpt_ne Full combined name (all 11 source datasets joined): see title above.Table of Contents
- Overview
- Dataset at a Glance
- Source Datasets That Were Merged
- File Format
- Top-Level Schema (Fields)
- `conversations` Structure (ShareGPT Format)
- `metadata_json` Sub-Fields
- Domain / Category / Sub-Domain Breakdown
- Behavior Diversity
- Question Pattern Diversity
- Sample Question–Answer Pairs (One Per Source)
- License Breakdown
- Text Length Statistics
- Coverage: Provinces & Fiscal Years
- How to Load the Dataset (Code Demos)
- Suggested Use Cases
- Known Limitations
- Credits / Attribution
- Comparison With the Companion Dataset (1,752-Row Merge)
1. Overview
This dataset is a merged collection of 11 separate Nepali-language health-statistics datasets, unified into a single .jsonl file using the ShareGPT conversational format ({"conversations": [{"from": "human", ...}, {"from": "gpt", ...}]}).
Every row is a single-turn, fact-based Question & Answer pair written entirely in Nepali (Devanagari script), sourced from real Nepal government health publications published by the Department of Health Services (DoHS). Each answer restates the exact factual figure, percentage, or rate asked about, making the dataset suitable for training or evaluating models on grounded, factual, Nepali-language question answering in the public-health domain — with a strong focus on maternal, adolescent, and reproductive health, immunization, and disease trend indicators.
The file that was uploaded and analyzed for this README is:
merged_new_health_332_sharegpt_ne.jsonl2. Dataset at a Glance
3. Source Datasets That Were Merged
The dataset name in the title of this README is the combination of the `source_name` value of every dataset merged into this file. The table below shows each individual source, how many rows it contributed, and its license.
All 11 sources share the same publisher (source_repo): Department of Health Services.
What each source actually covers:
4. File Format
- Format: JSON Lines (
.jsonl) — each line is one independent, valid JSON object (no surrounding array, no trailing commas between lines). - Encoding: UTF-8 (required, since almost all text is Devanagari script).
- Conversation schema: ShareGPT-style — a list of turns under the
conversationskey, alternatinghumanandgptroles. - Line count = record count: 332 lines = 332 records.
5. Top-Level Schema (Fields)
Every JSON object (row) in the file has exactly the following top-level fields:
⚠️ Important:metadata_jsonis stored as an escaped string, not a nested JSON object — you must runjson.loads()on it a second time after parsing the outer row.
6. conversations Structure (ShareGPT Format)
Each conversations value is a list of exactly 2 turns:
"conversations": [
{
"from": "human",
"value": "आर्थिक वर्ष २०७३/७४ मा सामान्य चिकित्सा बिरामी संख्या कति थियो?"
},
{
"from": "gpt",
"value": "आर्थिक वर्ष २०७३/७४ मा सामान्य चिकित्सा बिरामी संख्या ४१२०० थियो।"
}
]7. metadata_json Sub-Fields
Once parsed (json.loads(row["metadata_json"])), each metadata object contains the following keys:
Note: in this 332-row file, only `generation_sub_domain` varies (2 values) — every othermetadata_jsonfield, includinggeneration_category, is constant. This is a narrower, more topically-focused merge than the companion 1,752-row dataset.
8. Domain / Category / Sub-Domain Breakdown
8.1 generation_category (1 unique value — constant)
8.2 generation_sub_domain (2 unique values)
8.3 source_name ↔ generation_sub_domain Cross-Reference
8.4 id Prefix Breakdown
Unlike the companion 1,752-row dataset, this file's id prefixes map 1:1 onto source_name (no further topic-splitting occurs):
9. Behavior Diversity
Like its companion dataset, this file uses a structured metadata schema that formally labels every row with a behavior and behavior_definition.
9.1 Formal behavior Label (from metadata_json)
`behavior_definition` (constant across all rows):
➡️ Exactly as in the companion dataset, the formal metadata is intentionally uniform: 100% of rows are tagged with a single behavior class — "answer factually, in Nepali, without altering the underlying statistic." This is a single-behavior, high-precision factual QA dataset.
9.2 Practical Answer-Behavior Patterns Observed
Even though the metadata behavior tag never changes, the gpt answers exhibit several distinct, recurring answering behaviors depending on the underlying indicator type:
10. Question Pattern Diversity
Every human question in this dataset is unique as an exact string (332 out of 332 questions are exact-text-unique). To measure structural pattern diversity (i.e., how many distinct question templates exist once fiscal years, provinces, and numbers are treated as variables), all Devanagari/Arabic digits were stripped and the remaining unique phrasing was counted per source.
10.1 Pattern Diversity Per Source
Interpretation:patient_typesandnurses_regask a genuinely different question every single time (every row is a distinct patient/department type or a distinct registration category — 100% template diversity).immunizationandmalaria_trendreuse a moderate number of templates (24–31) across different vaccines/years. Sources likeadolescent_anc,fp_modern, andphc_orcreuse a small number of question templates (6–9) and vary mainly the province or fiscal year.
10.2 Example Question Templates — immunization (highest-diversity multi-row source)
10.3 Example Question Templates — Low-Diversity Sources
11. Sample Question–Answer Pairs (One Per Source)
12. License Breakdown
⚠️cc-by-sa(ShareAlike) is more restrictive than plaincc-by: any derivative dataset built substantially fromadolescent_abortion,adolescent_anc, orpatient_typesrows should be released under a compatible ShareAlike license.
13. Text Length Statistics
14. Coverage: Provinces & Fiscal Years
Fiscal-year mention frequency
This dataset's fiscal-year coverage is narrower and more recent than the companion 1,752-row dataset (which spans 9 fiscal years back to २०६५/६६) — over 64% of this file's fiscal-year-tagged questions concern २०७३/७४ (2016/17) alone.
15. How to Load the Dataset (Code Demos)
15.1 Plain Python (no dependencies)
import json
file_path = "merged_new_health_332_sharegpt_ne.jsonl"
records = []
with open(file_path, "r", encoding="utf-8") as f:
for line in f:
line = line.strip()
if not line:
continue
records.append(json.loads(line))
print(f"Total records loaded: {len(records)}")
# Inspect the first record
first = records[0]
print("ID:", first["id"])
print("Source:", first["source_name"])
question = first["conversations"][0]["value"]
answer = first["conversations"][1]["value"]
print("Q:", question)
print("A:", answer)
# metadata_json is a STRING — parse it separately
meta = json.loads(first["metadata_json"])
print("Sub-domain:", meta["generation_sub_domain"])15.2 Using pandas
import json
import pandas as pd
file_path = "merged_new_health_332_sharegpt_ne.jsonl"
# Read line-by-line (jsonl -> DataFrame)
rows = []
with open(file_path, "r", encoding="utf-8") as f:
for line in f:
line = line.strip()
if line:
rows.append(json.loads(line))
df = pd.DataFrame(rows)
# Expand the stringified metadata_json into real columns
meta_df = pd.json_normalize(df["metadata_json"].apply(json.loads))
df = pd.concat([df.drop(columns=["metadata_json"]), meta_df], axis=1)
# Pull question / answer text out of the conversations column
df["question"] = df["conversations"].apply(lambda c: c[0]["value"])
df["answer"] = df["conversations"].apply(lambda c: c[1]["value"])
print(df.shape)
print(df[["id", "source_name", "generation_sub_domain", "question", "answer"]].head())
# Example: filter to a single source
malaria = df[df["source_name"] == "malaria_trend"]
print(f"{len(malaria)} malaria-trend rows")15.3 Using pandas' built-in JSON-lines reader (shortcut)
import pandas as pd
df = pd.read_json("merged_new_health_332_sharegpt_ne.jsonl", lines=True)
print(df.shape)
print(df.columns.tolist())15.4 Using the Hugging Face datasets library
from datasets import load_dataset
dataset = load_dataset(
"json",
data_files="merged_new_health_332_sharegpt_ne.jsonl",
split="train",
)
print(dataset)
print(dataset[0])
# Filter by source_name
immunization_only = dataset.filter(lambda row: row["source_name"] == "immunization")
print(len(immunization_only))15.5 Streaming line-by-line (for very large files / low memory)
import json
def iter_records(path):
with open(path, "r", encoding="utf-8") as f:
for line in f:
line = line.strip()
if line:
yield json.loads(line)
for i, rec in enumerate(iter_records("merged_new_health_332_sharegpt_ne.jsonl")):
q = rec["conversations"][0]["value"]
a = rec["conversations"][1]["value"]
# process record here
if i < 3:
print(f"[{rec['id']}] Q: {q}\n[{rec['id']}] A: {a}\n")15.6 Converting to OpenAI-style / Alpaca-style fine-tuning format
import json
def to_openai_chat_format(record):
q = record["conversations"][0]["value"]
a = record["conversations"][1]["value"]
return {
"messages": [
{"role": "user", "content": q},
{"role": "assistant", "content": a},
]
}
with open("merged_new_health_332_sharegpt_ne.jsonl", "r", encoding="utf-8") as f_in, \
open("openai_format_332.jsonl", "w", encoding="utf-8") as f_out:
for line in f_in:
rec = json.loads(line)
f_out.write(json.dumps(to_openai_chat_format(rec), ensure_ascii=False) + "\n")15.7 Merging this file with the companion 1,752-row dataset
import json
files = [
"merged_health_dataset_sharegpt_ne__1_.jsonl",
"merged_new_health_332_sharegpt_ne.jsonl",
]
combined = []
seen_ids = set()
for path in files:
with open(path, "r", encoding="utf-8") as f:
for line in f:
line = line.strip()
if not line:
continue
rec = json.loads(line)
if rec["id"] in seen_ids:
raise ValueError(f"Duplicate id across files: {rec['id']}")
seen_ids.add(rec["id"])
combined.append(rec)
print(f"Combined total: {len(combined)} rows")
with open("combined_health_dataset_2084.jsonl", "w", encoding="utf-8") as f_out:
for rec in combined:
f_out.write(json.dumps(rec, ensure_ascii=False) + "\n")16. Suggested Use Cases
17. Known Limitations
- Single-turn only: No multi-turn or follow-up-question conversations are present (100% of rows have exactly 2 turns).
- Uniform behavior/category labels: The
behavior,question_type, andgeneration_categorymetadata fields carry no diversity in this file (always the same value) — genuine diversity exists only ingeneration_sub_domain(2 values) and in the underlying question phrasing/topic. - Template reuse in several sources: Sources such as
adolescent_anc,fp_modern, andphc_orcreuse a small number of question templates (6–9) and vary mainly the province or fiscal year being asked about — see Section 10. - Small sources:
nurses_reg(6 rows) andpatient_types(10 rows) are very small compared tomalaria_trend(68 rows), which may cause class imbalance if this dataset is used for per-source stratified evaluation. - Mixed licensing: 18.4% of rows (
adolescent_abortion,adolescent_anc,patient_types) are licensedcc-by-sa(ShareAlike) rather than plaincc-by. - No English content: The dataset is 100% Nepali; it is not suitable as-is for multilingual or English-Nepali parallel-text tasks.
- Narrow fiscal-year range: Only 3 fiscal years are referenced (२०७१/७२–२०७३/७४, roughly 2014–2017 Gregorian), a much narrower time span than the companion dataset's 9-year range.
18. Credits / Attribution
- Primary publisher: Department of Health Services (DoHS), Nepal — 100% of records in this file.
- Format: Compiled and merged into ShareGPT conversational JSONL format for instruction-tuning / QA use.
If you use this dataset, please retain attribution to the original Department of Health Services publications listed in the source field of each record.
19. Comparison With the Companion Dataset (1,752-Row Merge)
This dataset is designed to be schema-compatible with the earlier, larger merge (merged_health_dataset_sharegpt_ne__1_.jsonl). Key differences:
Both files use the identical top-level schema and `metadata_json` sub-schema, so they can be safely concatenated (see the merge code snippet in Section 15.7) as long as id uniqueness across both files is verified first.
