CoolFace
Datasetpublic

sabin1234/Major_Health_Indicator_FY_2073_74_Nepali_Statistical_QA

Major Health Indicators FY 2073/74 — Nepali Statistical Q&A Dataset Overview This dataset (nepali_sharegpt_suddha_nepali_final_cleaned.jsonl) is a collection of 587 instruction-following conversation pairs in Nepali, built entirely from Nepal's official Department of Health Services (DoHS) "Major Health Indicators" report for fiscal year 2073/74 (2016/17 AD). Each record is a single-turn human↔gpt exchange: a Nepali-language question asking for a specific… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Major_Health_Indicator_FY_2073_74_Nepali_Statistical_QA.

sourceHugging Facecc-by-4.0updated 18d agoView on Hugging Face
0likes53downloads
Dataset Card

Major Health Indicators FY 2073/74 — Nepali Statistical Q&A Dataset

Overview

This dataset (nepali_sharegpt_suddha_nepali_final_cleaned.jsonl) is a collection of 587 instruction-following conversation pairs in Nepali, built entirely from Nepal's official Department of Health Services (DoHS) "Major Health Indicators" report for fiscal year 2073/74 (2016/17 AD). Each record is a single-turn human↔gpt exchange: a Nepali-language question asking for a specific health-infrastructure statistic — at the national, provincial, or district level — followed by a short, data-grounded Nepali-language answer stating the exact figure.

This dataset is structurally and topically very different from the other FAQ-style datasets in this series: instead of open-ended explanatory questions, it is a large set of templated statistical lookup questions (e.g., "How many government hospitals were there in Province 1 in FY 2073/74?") systematically repeated across Nepal's administrative levels (national → 7 provinces → 77 districts) for each of 7 health-system indicators.

  • —File format: JSON Lines (.jsonl), one JSON object per line
  • —Total records: 587
  • —Source: Department of Health Services — Major Health Indicators FY 2073/74 report
  • —License: CC-BY-4.0 (permissive tier) — note this differs from the Apache-2.0 license used by every other dataset in this series
  • —Task type: instruction-following
  • —Generation type / condition: synthetic (questions were synthetically generated from the underlying statistical report, unlike the "real"/"scraped" FAQ transcriptions in sibling datasets)
  • —Source revision: dohs_2073_74

⚠️ Important nuance: metadata says "English," but the content is Nepali

Every record's provenance metadata reports language: en, script: Latn, content_language: English, content_script: Latin, and english_content_allowed: true — but the actual question/answer text in every single record (587/587) is written in Devanagari-script Nepali, with zero Latin-script mixing found anywhere in the conversation text. This indicates the metadata fields still describe the original English-language source dataset (visible in the embedded Windows file path major_health_indicators_fy2073_74_sharegpt_en_open_clean.jsonl) from before it was translated/localized into "suddha" (pure) Nepali — the metadata was not updated to reflect the translated content. This is the only dataset in the series where the declared content-language metadata does not match the actual language of the text.

File / Record Structure

This dataset uses a different, nested top-level structure compared to the other datasets in the series — provenance fields are wrapped inside a source_provenance object rather than sitting at the top level:

FieldDescriptionObserved value(s)
idUnique record identifiersg_<32-char hex hash>, e.g. sg_cef406f29e207b6fb383064f27228928 (all 587 unique)
conversationsList of 2 turns: human (question) then gpt (answer)always exactly 2 turns
input_formatFormat labelsharegpt (constant)
source_task_idTask identifieridentical to id for every record
source_provenanceNested object containing all source/licensing/metadata fields (see below)—

source_provenance sub-fields

FieldDescriptionObserved value(s)
idSame as top-level id—
sourceFull source descriptorDepartmentOfHealthServices/Major_Health_Indicators_FY2073_74:default:train (constant)
source_nameInternal dataset namemajor_health_indicators_fy2073_74 (constant)
source_repoSource repositoryDepartmentOfHealthServices/Major_Health_Indicators_FY2073_74 (constant)
source_configConfig/subset namedefault (constant)
source_splitData splittrain (constant)
source_revisionSource version tagdohs_2073_74 (constant)
source_row_idRow ID in the original sourceid with a :1 suffix appended
language / language_code / scriptDeclared as English/eng/Latin (see nuance above)en / — / Latn (constant, but does not reflect actual content)
licenseLicenseCC-BY-4.0 (constant)
license_tierLicense permissiveness tierpermissive (constant)
task_typeTask categoryinstruction-following (constant)
generation_typeHow the data was producedsynthetic (constant)
conditionData condition/type flagsynthetic (constant)
urlSource URLempty string (constant)
metadata_jsonStringified JSON with generation-level metadata (see below)—
dataset_nameLocal file path of the pre-translation source filea Windows path ending in major_health_indicators_fy2073_74_sharegpt_en_open_clean.jsonl (constant across all records)
dataset_splitSplit labeltrain (constant)
row_indexRow index within the source filevaries (0-based, sequential per sub-domain group)

metadata_json sub-fields (nested inside source_provenance)

FieldDescriptionObserved value
generation_domainBroad subject domainPublic Health (constant)
generation_categoryCategoryHealth Indicators FY 2073/74 (constant)
generation_sub_domainSpecific indicator type7 distinct values (see distribution below)
behaviorExpected model behaviorshort factual answer (constant)
behavior_definitionDetailed instruction"Provide a concise, data-grounded answer while preserving the key fact." (constant)
question_typeType of questiongrounded open-ended question (constant) — a different label from the "fact-based question" wording used elsewhere in the series
question_lengthTarget question length range"60 to 220 characters" (constant, expressed as a range string rather than a per-record number)
response_lengthTarget response length range"15 to 120 characters" (constant, range string)
maximum_response_sentencesCap on sentences allowed in the answer2 (constant)
minimum_reasoning_dimensionsMinimum reasoning dimensions required1 (constant)
minimum_complexity_scoreMinimum complexity score required4 (constant)
content_languageDeclared content languageEnglish (constant — see nuance above)
content_scriptDeclared content scriptLatin (constant — see nuance above)
english_content_allowedWhether English content is permittedtrue (constant — the only dataset in the series where this is true)

Distribution Analysis

Sub-domain Distribution (the dataset's main axis of variation)

This is a single-domain, single-category dataset (Public Health → Health Indicators FY 2073/74), but it spans 7 distinct health-infrastructure/service indicators, each queried at multiple administrative levels:

Sub-domainCount%What it covers
GoN Hospital8514.5%Number of government (GoN) hospitals
PHC/ORC8514.5%Number of Primary Health Care Outreach Clinics
EPI Clinics8514.5%Number of Expanded Programme on Immunization clinics
FCHVs8514.5%Number of Female Community Health Volunteers
Health Post8414.3%Number of health posts (स्वास्थ्य चौकी)
BCG8414.3%BCG vaccine coverage percentage
PHCC7913.5%Number of Primary Health Care Centers

Each sub-domain contains one question at the national level, one for each of Nepal's 7 provinces, and one for most/all of Nepal's 77 districts — i.e., roughly 1 (national) + 7 (provincial) + 77 (district) = 85 questions per indicator for most sub-domains. PHCC (79) and the two 84-count sub-domains (Health Post, BCG) have slightly fewer, likely because PHCC/Health Post/BCG data was not available or not applicable for every district in the original report.

Behavior / Question Type / License Distribution

FieldValueCount%
behaviorshort factual answer587100%
question_typegrounded open-ended question587100%
licenseCC-BY-4.0587100%
generation_type / conditionsynthetic587100%

All other categorical fields listed in the tables above are likewise 100% constant across all 587 records — this dataset has no variation in domain, category, behavior, question type, generation constraints, or licensing; the only variation is in generation_sub_domain (7 values) and, of course, the specific administrative unit and number named in each question/answer.

Conversation Structure

MetricValue
Turns per conversationAlways 2 (1 human + 1 gpt)
Roles usedhuman, gpt
Human turns587
GPT turns587

Single-turn instruction/response format, consistent with the rest of the series.

Text Length Statistics

MetricMinMaxAverage
Question length (characters)55103~70
Answer length (characters)55101~69

Both questions and answers are tightly clustered in length — this is expected given the templated, formulaic nature of the questions (e.g., "आर्थिक वर्ष २०७३/७४ मा [प्रदेश/जिल्ला] मा कतिवटा [सुविधा प्रकार] थिए?" — "How many [facility type] were there in [province/district] in FY 2073/74?") and answers (a single sentence restating the question with the specific number filled in). This makes it by far the most length-uniform dataset in the series, and the shortest average answer length overall (~69 characters, well below even the SAC Nepal fitness dataset's ~121 characters).

Sample Records

Sub-domainExample QuestionExample Answer
GoN Hospitalआर्थिक वर्ष २०७३/७४ मा राष्ट्रियमा कतिवटा सरकारी अस्पतालहरू थिए?आर्थिक वर्ष २०७३/७४ मा राष्ट्रियमा १२३ वटा सरकारी अस्पतालहरू थिए।
PHCCआर्थिक वर्ष २०७३/७४ मा प्रदेश १ मा कति वटा प्राथमिक स्वास्थ्य केन्द्रहरू थिए?आर्थिक वर्ष २०७३/७४ मा प्रदेश १ मा ४१ वटा प्राथमिक स्वास्थ्य केन्द्रहरू थिए।
Health Postआर्थिक वर्ष २०७३/७४ मा ०१ ताप्लेजुङमा कति वटा स्वास्थ्य चौकीहरू थिए?आर्थिक वर्ष २०७३/७४ मा ०१ ताप्लेजुङमा ५० वटा स्वास्थ्य चौकीहरू थिए।
FCHVsआर्थिक वर्ष २०७३/७४ मा राष्ट्रिय स्तरमा कुल कति जना महिला सामुदायिक स्वास्थ्य स्वयंसेविकाहरू रहेका थिए?आर्थिक वर्ष २०७३/७४ मा राष्ट्रिय स्तरमा ४९००१ जना महिला सामुदायिक स्वास्थ्य स्वयंसेविकाहरू रहेका थिए।
BCGआर्थिक वर्ष २०७३/७४ मा राष्ट्रिय स्तरमा बीसीजी खोपको कभरेज कति थियो?आर्थिक वर्ष २०७३/७४ मा राष्ट्रिय स्तरमा बीसीजी खोपको कभरेज ९१.३% थियो।

Summary

This dataset is fundamentally different in character from the other Nepali health FAQ datasets in this series: rather than explanatory FAQ content, it is a structured, templated statistical Q&A set distilled from a single official government report (DoHS Major Health Indicators, FY 2073/74). It systematically asks the same 7 kinds of infrastructure/coverage questions (hospital counts, PHCC counts, health post counts, PHC/ORC clinic counts, EPI clinic counts, FCHV volunteer counts, and BCG vaccination coverage) across all administrative levels of Nepal — national, all 7 provinces, and nearly all 77 districts — for a total of 587 records.

Key distinguishing features versus the rest of the series:

  • —License: CC-BY-4.0, not Apache-2.0
  • —Generation type: synthetic (templated from statistics), not real/scraped
  • —Question type: "grounded open-ended question," not "fact-based question"
  • —Metadata/content language mismatch: metadata declares English/Latin, but all actual text is Nepali/Devanagari
  • —Record structure: nested under source_provenance, unlike the flat structure used elsewhere
  • —Content style: short, numeric, single-fact lookups rather than explanatory prose — by far the shortest and most uniform answers in the series (~69 characters average)

This kind of dataset is well suited for fine-tuning or evaluating a language model's ability to retrieve and state precise government health statistics in Nepali, especially for administrative/geographic-level lookups, and could complement the more explanatory FAQ-style datasets in this series by adding a grounded-numeric-fact dimension to a broader Nepali health instruction-tuning corpus.