CoolFace
Datasetpublic

sabin1234/Nepal_Education_Provincial_Budget_Tax_Miscellaneous_Statistics

Nepal Education, Provincial Budget, Tax & Miscellaneous Statistics — Instruction Dataset A Nepali-language (with a small English subset) instruction-following (Q&A) dataset built from official Nepali government and statistical sources — covering federal/provincial education budgets, tax exemptions, trade statistics, cooperative membership, disaster losses, national accounts, and Department of Roads budget data. File: merged_all_edu_prov_tax_misc_serial.jsonl Format: JSON Lines… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Nepal_Education_Provincial_Budget_Tax_Miscellaneous_Statistics.

sourceHugging Facecc-by-4.0updated 16d agoView on Hugging Face
0likes45downloads
Dataset Card

Nepal Education, Provincial Budget, Tax & Miscellaneous Statistics — Instruction Dataset

A Nepali-language (with a small English subset) instruction-following (Q&A) dataset built from official Nepali government and statistical sources — covering federal/provincial education budgets, tax exemptions, trade statistics, cooperative membership, disaster losses, national accounts, and Department of Roads budget data.

File: merged_all_edu_prov_tax_misc_serial.jsonl Format: JSON Lines (.jsonl) — one JSON object per line Total records: 1,680


1. Overview

PropertyValue
File namemerged_all_edu_prov_tax_misc_serial.jsonl
FormatJSONL (newline-delimited JSON)
Total rows / conversations1,680
Turns per conversation2 (1 human turn + 1 gpt turn) — always single-turn Q&A
Total human turns1,680
Total gpt turns1,680
Primary languageNepali (ne / ISO 639-3 npi), Devanagari script
Secondary languageEnglish (en), Latin script — 51 rows
Task typeinstruction-following (100%)
Generation typereal (100% — answers are derived from real statistical source tables, not synthetic/hallucinated)
Data domainsEducation budgets, Provincial budgets & indicators, Tax exemptions/concessions, National accounts, Financial inclusion (cooperatives), Trade & customs, Disaster & risk, Public infrastructure finance
Source documents9 distinct government/statistical publications
LicensesCC0 (Public Domain) and CC-BY, both permissive

2. Record Schema

Each line is a single JSON object with the following top-level fields:

FieldTypeDescription
idstringUnique record ID, encodes the "question pattern family" + a serial number (e.g. edu_budget_faq_0007)
conversationsarray of objectsThe Q&A pair: [{"from": "human", "value": "..."}, {"from": "gpt", "value": "..."}]
sourcestringHuman-readable title of the original source publication/dataset
source_namestring \nullShort machine-friendly slug for the source (null for the Department of Roads subset)
source_repostringPublishing organization / repository the data came from
source_configstring \nullDataset config label (default or null)
source_splitstringAlways train
source_revisionstringAlways v1
source_row_idstringSame value as id (row-level traceability)
languagestringISO 639-1 code: ne or en
language_codestringISO 639-3 code: npi or en
scriptstringWriting script: Deva (Devanagari) or Latn (Latin)
licensestring \nullCC0: Public Domain, cc-by, or null
license_tierstring \nullpermissive or null
task_typestringAlways instruction-following
generation_typestringAlways real
conditionstringAlways real
urlstring \nullSource URL (empty/blank for all current rows)
metadata_jsonstring (JSON-encoded)Nested metadata object — see Section 4

conversations structure

json
"conversations": [
  {"from": "human", "value": "<question in Nepali or English>"},
  {"from": "gpt",   "value": "<factual answer, same language as the question>"}
]

Every record has exactly 2 turns: one human question and one gpt answer — this is a single-turn factual QA dataset, not multi-turn dialogue.

metadata_json structure (parsed)

metadata_json is a string containing a nested JSON object. Once parsed, it contains:

Sub-fieldDescription
generation_domainBroad topical domain (e.g. "शिक्षा बजेट" / "National Accounts")
generation_categoryNarrower category within the domain
generation_sub_domainFiner-grained sub-topic (e.g. specific year, region type)
behaviorThe intended model behavior for this Q&A pair (see Section 6)
behavior_definitionA one-sentence Nepali description of what "correct" behavior means for this row
question_typeHigh-level question category (fact-based)
content_languageHuman-readable language name ("नेपाली" / "English")
content_scriptHuman-readable script name ("देवनागरी" / "Latin")
english_content_allowedBoolean — whether English terms/numerals are permitted in the answer

3. Source Publications (9 total)

All answers are grounded in real, published Nepali statistical/government data — no hallucinated facts.

Source (full title)`source_name`Publisher (`source_repo`)RecordsLicense
Analysis of Provincial Budgets and Education Indicators FY 2021/22provincial_budget_educationAgriculture Inputs Company Limited / CBS488CC0
Analysis of Federal Education Budget FY 2021/22federal_education_budgetCentral Bureau of Statistics317CC0
Intermediate Consumption by Industrial Division in India 2000–2018 at Current Prices (Rs Millions)intermediate_consumption_statsCentral Bureau of Statistics288cc-by
Nepal Tax Data on Exemptions and Concessionstax_exemptions_concessionsMinistry of Finance208CC0
Households with Membership of Co-operatives and Saving Groupscoop_membership_statsNepal Household Survey155CC0
Imports and Exports by Nepal Customs Offices for First Month of FY 2074/75customs_import_export_statsNepal Customs Trade Data80CC0
Loss of Lives and Properties from Disastersdisaster_losses_statsMinistry of Finance Disaster Data56CC0
सडक विभाग (Department of Roads budget records, FY 2064/65)(null)सडक विभाग (Department of Roads)51(null / unspecified)
Analysis of Federal Education Budget FY 2021/22 — Government Units Shareeducation_budget_share_by_unitCentral Bureau of Statistics37cc-by
Note: The "Intermediate Consumption" source title mentions "India" but the questions and answers themselves are about Nepali economic sectors (कृषि तथा वन, etc.) — the title appears to be inherited as-is from the upstream source metadata.

4. Domain / Category / Sub-domain Breakdown

4.1 generation_domain (8 unique values)

DomainRecords% of dataset
प्रादेशिक बजेट तथा शिक्षा (Provincial Budget & Education)48829.0%
शिक्षा बजेट (Education Budget)35421.1%
National Accounts28817.1%
कर तथा राजस्व (Tax & Revenue)20812.4%
Financial Inclusion1559.2%
Trade & Customs804.8%
Disaster & Risk563.3%
public infrastructure finance513.0%

4.2 generation_category (9 unique values)

CategoryRecords
प्रादेशिक बजेट र शिक्षा सूचक488
संघीय शिक्षा बजेट विश्लेषण317
Intermediate Consumption by Industry288
कर छुट तथा रियायत208
Co-operative Membership Statistics155
Import Export Statistics80
Loss of Lives and Properties56
Department of Roads budget and expenditure51
सरकारी तहअनुसार शिक्षा बजेट अंश37

4.3 generation_sub_domain (17 unique values)

Sub-domainRecords
प्रदेशस्तरीय तथ्यांक488
Industrial Division288
बजेट तथ्यांक258
कर खर्च208
Customs Trade80
बजेट तथा खर्च59
Disaster Losses56
Fiscal Year 2064/6551
Eco-Development Region48
संघ प्रदेश स्थानीय37
Additional Unique35
Comparative30
NAPA Combined Vulnerability Index10
Bio-Climatic Zone10
Ecological Belt9
Urban/Rural8
National5

5. Language, Script & License Distribution

5.1 Language

LanguageCode (ISO 639-1)Code (ISO 639-3)ScriptRecords%
NepalinenpiDeva (Devanagari)1,62997.0%
EnglishenenLatn (Latin)513.0%

5.2 License

LicenseLicense tierRecords%
CC0: Public Domainpermissive1,39483.0%
cc-bypermissive23514.0%
(unspecified / null)(null)513.0%

5.3 Task / Generation metadata (constant fields)

FieldValueCoverage
task_typeinstruction-following100% (1,680/1,680)
generation_typereal100% (1,680/1,680)
conditionreal100% (1,680/1,680)
source_splittrain100% (1,680/1,680)
source_revisionv1100% (1,680/1,680)
english_content_allowedfalse96.9% (1,629/1,680)
english_content_allowedtrue3.1% (51/1,680)

6. Behavior Diversity

The behavior and behavior_definition metadata fields describe what the model is expected to do when answering each question. There are 3 distinct behavior types, each paired with a more detailed behavior definition (4 unique full definitions total, since one behavior has two slightly different phrasings across subsets).

`behavior` (metadata)RecordsAssociated `behavior_definition` (Nepali)English meaning
तथ्यमा आधारित उत्तर दिने (Give a fact-based answer)1,050"मूल तथ्य कायम राख्दै शुद्ध नेपालीमा उत्तर दिने।"Preserve the source fact exactly, and answer in pure/correct Nepali.
Provide factual answer derived only from the dataset579"मूल प्रश्नको अर्थ र मूल उत्तरको तथ्य कायम राख्दै शुद्ध नेपालीमा रूपान्तरण गरिएको उत्तर दिने।"Preserve both the meaning of the original question and the fact of the original answer, rendering the answer as a faithful Nepali translation/adaptation.
provide precise factual financial answers51"उपलब्ध सडक कार्यालयका बजेट तथा खर्च तथ्याङ्कका आधारमा सिधा र तथ्यपरक उत्तर दिने।"Give a direct, fact-based answer strictly grounded in the available Road Office budget/expenditure data.

A closely related variant of the second behavior definition also appears on a smaller subset:

`behavior_definition` variantRecords
मूल तथ्य कायम राख्दै शुद्ध नेपालीमा उत्तर दिने।954
मूल प्रश्नको अर्थ र मूल उत्तरको तथ्य कायम राख्दै शुद्ध नेपालीमा रूपान्तरण गरिएको उत्तर दिने।579
मूल प्रश्नको अर्थ र मूल उत्तरको तथ्य कायम राख्दै शुद्ध नेपालीमा उत्तर दिने।96
उपलब्ध सडक कार्यालयका बजेट तथा खर्च तथ्याङ्कका आधारमा सिधा र तथ्यपरक उत्तर दिने।51

Common thread across all behaviors: every answer must be (1) strictly grounded in the source statistical table (no fabrication), and (2) delivered in fluent, natural Nepali (or English, for the 51 English rows) rather than a literal/awkward translation.

question_type (2 unique values)

`question_type`Records%
तथ्यमा आधारित प्रश्न (fact-based question)1,10165.5%
fact-based57934.5%

(Both labels describe the same underlying question type — "fact-based" — just recorded in Nepali vs. English metadata for different subsets of the merge.)


7. Question Pattern Diversity

Although every question is "fact-based," the dataset contains 34 distinct question-pattern families, identifiable via the id prefix (the part of the id before the trailing serial number, e.g. edu_budget_faq_0007 → family edu_budget_faq). Each family represents a repeatable question template applied across different years, provinces, ministries, tax categories, etc. Two example Q&A pairs are shown per family below.

#ID Prefix (pattern family)RecordsExample QuestionExample Answer
1inter_cons_faq288कृषि तथा वन को २०००/०१ मा चालु मूल्यमा मध्यवर्ती उपभोग कति थियो?कृषि तथा वन को २०००/०१ मा मध्यवर्ती उपभोग रु. ५५,२६८ मिलियन थियो।
2coop_membership_faq155सहरी क्षेत्रका कति प्रतिशत घरपरिवार सहकारी तथा बचत समूहका सदस्य छन्?सहरी क्षेत्रका ५३.५७ प्रतिशत घरपरिवार सहकारी तथा बचत समूहका सदस्य छन्।
3prov_edu_ind147कोशी प्रदेश मा विद्यालय संख्या कति थियो?कोशी प्रदेश मा विद्यालय संख्या ६९५८ थियो।
4customs_trade_faq80आर्थिक वर्ष २०७४/७५ को पहिलो महिनामा वीरगञ्ज भन्सार कार्यालयले कति आयात मूल्य व्यवस्थापन गर्‍यो?वीरगञ्ज भन्सार कार्यालयले रु. २६,१९५,०७१.९३ हजार आयात मूल्य अभिलेख गर्‍यो।
5tax_income63२०२४ मा कर दायराभित्र घोषित कृषि व्यवसाय को कर खर्च कति थियो?२०२४ मा कर दायराभित्र घोषित कृषि व्यवसाय को कर खर्च ३२८८७ थियो।
6edu_budget_faq59शिक्षा विज्ञान तथा प्रविधि मन्त्रालय को लागि आ.व. २०१९/२० को बजेट कति छुट्याइएको थियो?शिक्षा विज्ञान तथा प्रविधि मन्त्रालय को लागि आ.व. २०१९/२० मा २३.८९८७ बजेट छुट्याइएको थियो।
7prov_bud_exp58कोशी प्रदेश को बजेट कति थियो?कोशी प्रदेश को बजेट ४०.९ थियो।
8disaster_loss_faq56कुन प्रकारको प्रकोपले सबैभन्दा बढी घटना संख्या अभिलेख गर्‍यो?चट्याङले १७७ घटनासहित सबैभन्दा बढी घटना संख्या अभिलेख गर्‍यो।
9edu_expenditure_faq51२०१५/१६ मा कुल बजेट कति थियो?२०१५/१६ मा कुल बजेट ८१९.१६ थियो।
10prov_revenue51कोशी प्रदेश को संघीय सरकार अनुदान कति थियो?कोशी प्रदेश को संघीय सरकार अनुदान १५.२९५७ थियो।
11dor_budget_2064_6551बाँडफाँट नगरिएको बजेट रकम कति थियो?बाँडफाँट नगरिएको बजेट रकम ३,९६५,६०० थियो।
12tax_consumption46२०२४ मा कृषि वन माछापालन खाद्य पेय सुर्ती को अन्तिम उपभोग कति थियो?२०२४ मा कृषि वन माछापालन खाद्य पेय सुर्ती को अन्तिम उपभोग १०९१३५७४४ थियो।
13prov_responsive44कोशी प्रदेश को प्रत्यक्ष सहयोगी कति थियो?कोशी प्रदेश को प्रत्यक्ष सहयोगी २९.६२ थियो।
14edu_sector_unit_faq43संघ तहमा शिक्षा को प्रतिशत कति थियो?संघ तहमा शिक्षा को प्रतिशत ३९.८७ थियो।
15prov_edu_detail43कोशी प्रदेश को २०२१/२२ शिक्षा बजेट कति थियो?कोशी प्रदेश को २०२१/२२ शिक्षा बजेट १.१९०८ थियो।
16prov_gdp_hdi42कोशी प्रदेश को राष्ट्रिय जीडीपीमा अंश कति प्रतिशत थियो?कोशी प्रदेश को राष्ट्रिय जीडीपीमा अंश १५.८२ प्रतिशत थियो।
17tax_desc40२०२४ मा आधारभूत कृषि उत्पादन को कर खर्च कति थियो?२०२४ मा आधारभूत कृषि उत्पादन को कर खर्च ९०२१४७३७ थियो।
18edu_share_faq37२०२१/२२ मा संघ तहको शिक्षा बजेट अंश कति थियो?२०२१/२२ मा संघ तहको शिक्षा बजेट अंश ६०.१०६७ थियो।
19prov_rec_cap30कोशी प्रदेश को चालु कति थियो?कोशी प्रदेश को चालु १४.१६ थियो।
20edu_growth_infl_faq27२०१५/१६ को पूर्वानुमानित मुद्रास्फीति दर कति थियो?२०१५/१६ को पूर्वानुमानित मुद्रास्फीति दर ७ थियो।
21edu_grant_faq26प्रदेश को वित्तीय समानीकरण अनुदान कति थियो?प्रदेश को वित्तीय समानीकरण अनुदान ५७.९५४८ थियो।
22responsive_budget_faq26२०२१/२२ मा प्रत्यक्ष सहयोगी बजेटको प्रतिशत कति थियो?२०२१/२२ मा प्रत्यक्ष सहयोगी बजेटको प्रतिशत ६७.७६ थियो।
23edu_total_budget_faq26२०१४-१५ मा शिक्षा बजेट कति थियो?२०१४-१५ मा शिक्षा बजेट ८६.०३ अर्ब थियो।
24prov_local_exp26कोशी प्रदेश को कुल खर्च कति थियो?कोशी प्रदेश को कुल खर्च ८४.५२८९ थियो।
25prov_gdp_growth24२०१९/२० मा कोशी प्रदेश को जीडीपी वृद्धि दर कति थियो?२०१९/२० मा कोशी प्रदेश को जीडीपी वृद्धि दर ७.०५ थियो।
26prov_edu_budget23कोशी प्रदेश को कुल बजेट कति थियो?कोशी प्रदेश को कुल बजेट ३२.४६९२ थियो।
27edu_level_faq20२०२१/२२ मा पूर्व प्राथमिक तथा प्राथमिक शिक्षा को बजेट कति थियो?२०२१/२२ मा पूर्व प्राथमिक तथा प्राथमिक शिक्षा को बजेट ६०.४४४६ थियो।
28edu_program_faq19२०२१/२२ मा राष्ट्रपति शैक्षिक सुधार कोष को बजेट कति थियो?२०२१/२२ मा राष्ट्रपति शैक्षिक सुधार कोष को बजेट १० थियो।
29tax_vs_gdp18२०२४ मा मूल्य अभिवृद्धि कर खर्च जीडीपीको कति प्रतिशत थियो?२०२४ मा मूल्य अभिवृद्धि कर खर्च जीडीपीको ३.१५ प्रतिशत थियो।
30tax_concession16२०२४ मा रियायत को प्रतिशत कति थियो?२०२४ मा रियायत को प्रतिशत ९२.८२ थियो।
31tax_vat_head14२०२४ मा मूल्य अभिवृद्धि कर छुट कारोबार को कर खर्च कति थियो?२०२४ मा मूल्य अभिवृद्धि कर छुट कारोबार को कर खर्च १३७७४४९८९ थियो।
32tax_domestic11२०२४ मा आन्तरिक उत्पादन को प्रतिशत कति थियो?२०२४ मा आन्तरिक उत्पादन को प्रतिशत ८६.५६ थियो।
33edu_budget_source_faq10२०२१/२२ मा आन्तरिक राजस्व बाट कति रकम प्राप्त भएको थियो?२०२१/२२ मा आन्तरिक राजस्व बाट १०२४.९०७ रकम प्राप्त भएको थियो।
34edu_revenue_struct_faq10संघको राजस्व स्रोत के हो?संघको राजस्व स्रोत राजस्व हो।

Total across all 34 families: 1,680 (matches total record count exactly).

7.1 Recurring question patterns observed across families

Pattern typeExample phrasingWhere it appears
"How much was X in year Y?""...कति थियो?"Almost every family (budget, revenue, GDP, tax figures)
"What percentage was X?""...को प्रतिशत कति थियो?"tax_*, edu_share_faq, coop_membership_faq, prov_gdp_hdi
"Which X had the highest/largest Y?" (superlative)"कुन प्रकारको प्रकोपले सबैभन्दा बढी..."disaster_loss_faq, customs_trade_faq
"What is the definition/source of X?""...को राजस्व स्रोत के हो?"edu_revenue_struct_faq
Province-specific numeric lookup"कोशी प्रदेश को ... कति थियो?" (repeated per province)All prov_* families
Region/category comparative lookup"...तथा वन को ... मध्यवर्ती उपभोग कति थियो?" (repeated per industry & year)inter_cons_faq
Ministry/program-specific budget lookup"...मन्त्रालय...को लागि आ.व. ... बजेट कति छुट्याइएको थियो?"edu_budget_faq, edu_program_faq
Count-of-events lookup"डाटासेटमा बाढीका कति घटना अभिलेख गरिएका छन्?"disaster_loss_faq

8. Text Length Statistics

MetricMinMaxAverage
Question length (characters)2512057.8
Answer length (characters)2216260.6

All Q&A pairs are short, single-fact statements (typically 1 sentence each) — no multi-paragraph or multi-fact answers.


9. Sample Records

Example 1 — Federal education budget (Nepali):

json
{
  "id": "edu_budget_source_faq_0001",
  "conversations": [
    {"from": "human", "value": "२०२१/२२ मा आन्तरिक राजस्व बाट कति रकम प्राप्त भएको थियो?"},
    {"from": "gpt", "value": "२०२१/२२ मा आन्तरिक राजस्व बाट १०२४.९०७ रकम प्राप्त भएको थियो।"}
  ],
  "source": "Analysis of Federal Education Budget FY 2021/22",
  "source_name": "federal_education_budget",
  "source_repo": "Central Bureau of Statistics",
  "language": "ne",
  "language_code": "npi",
  "script": "Deva",
  "license": "CC0: Public Domain",
  "license_tier": "permissive",
  "task_type": "instruction-following",
  "generation_type": "real",
  "condition": "real",
  "metadata_json": "{\"generation_domain\": \"शिक्षा बजेट\", \"generation_category\": \"संघीय शिक्षा बजेट विश्लेषण\", \"generation_sub_domain\": \"बजेट तथ्यांक\", \"behavior\": \"तथ्यमा आधारित उत्तर दिने\", \"behavior_definition\": \"मूल तथ्य कायम राख्दै शुद्ध नेपालीमा उत्तर दिने।\", \"question_type\": \"तथ्यमा आधारित प्रश्न\", \"content_language\": \"नेपाली\", \"content_script\": \"देवनागरी\", \"english_content_allowed\": false}"
}

Example 2 — Department of Roads budget (the English/no-source_name subset):

json
{
  "conversations": [
    {"from": "human", "value": "बाँडफाँट नगरिएको बजेट रकम कति थियो?"},
    {"from": "gpt", "value": "बाँडफाँट नगरिएको बजेट रकम ३,९६५,६०० थियो।"}
  ],
  "source": "सडक विभाग",
  "source_name": null,
  "source_repo": "सडक विभाग",
  "language": "en",
  "script": "Latn",
  "license": null,
  "metadata_json": "{\"generation_domain\": \"public infrastructure finance\", \"generation_category\": \"Department of Roads budget and expenditure\", \"generation_sub_domain\": \"Fiscal Year 2064/65\", \"behavior\": \"provide precise factual financial answers\", \"behavior_definition\": \"उपलब्ध सडक कार्यालयका बजेट तथा खर्च तथ्याङ्कका आधारमा सिधा र तथ्यपरक उत्तर दिने।\", \"question_type\": \"तथ्यमा आधारित प्रश्न\", \"content_language\": \"English\", \"content_script\": \"Latin\", \"english_content_allowed\": true}"
}
Note: despite language/script being tagged en/Latn for this subset, the actual question/answer text is still in Devanagari — this reflects how the metadata was tagged upstream rather than the actual script of the content. Treat content_language / content_script inside metadata_json as the more reliable per-row indicator, and always verify against the actual conversations text if language purity matters for your use case.

10. How to Load This Dataset

10.1 Plain Python (standard library only)

python
import json

path = "merged_all_edu_prov_tax_misc_serial.jsonl"

records = []
with open(path, "r", encoding="utf-8") as f:
    for line in f:
        line = line.strip()
        if not line:
            continue
        records.append(json.loads(line))

print(f"Loaded {len(records)} records")

# Inspect the first record
first = records[0]
question = first["conversations"][0]["value"]
answer = first["conversations"][1]["value"]
print("Q:", question)
print("A:", answer)

# Parse the nested metadata
metadata = json.loads(first["metadata_json"])
print("Domain:", metadata["generation_domain"])
print("Behavior:", metadata["behavior"])

10.2 Using pandas

python
import json
import pandas as pd

path = "merged_all_edu_prov_tax_misc_serial.jsonl"

# jsonlines files load directly with lines=True
df = pd.read_json(path, lines=True)

# Expand the nested metadata_json string into real columns
meta_df = pd.json_normalize(df["metadata_json"].apply(json.loads))
df = pd.concat([df.drop(columns=["metadata_json"]), meta_df], axis=1)

# Extract question / answer as separate columns for convenience
df["question"] = df["conversations"].apply(lambda c: c[0]["value"])
df["answer"]   = df["conversations"].apply(lambda c: c[1]["value"])

print(df.shape)
print(df[["id", "question", "answer", "generation_domain", "behavior"]].head())

# Example: filter to only education-budget questions
edu_df = df[df["generation_domain"].str.contains("शिक्षा", na=False)]
print(f"Education-related rows: {len(edu_df)}")

10.3 Using the Hugging Face datasets library

python
from datasets import load_dataset

dataset = load_dataset(
    "json",
    data_files="merged_all_edu_prov_tax_misc_serial.jsonl",
    split="train"
)

print(dataset)
print(dataset[0])

# Filter by language
nepali_only = dataset.filter(lambda ex: ex["language"] == "ne")
english_only = dataset.filter(lambda ex: ex["language"] == "en")
print(f"Nepali rows: {len(nepali_only)}, English-tagged rows: {len(english_only)}")

# Map to extract question/answer as top-level fields
def extract_qa(example):
    example["question"] = example["conversations"][0]["value"]
    example["answer"] = example["conversations"][1]["value"]
    return example

dataset = dataset.map(extract_qa)

10.4 Streaming (for very large files / low-memory environments)

python
import json

def stream_records(path):
    with open(path, "r", encoding="utf-8") as f:
        for line in f:
            line = line.strip()
            if line:
                yield json.loads(line)

for record in stream_records("merged_all_edu_prov_tax_misc_serial.jsonl"):
    q = record["conversations"][0]["value"]
    a = record["conversations"][1]["value"]
    # process one record at a time without loading the whole file into memory

10.5 Converting to a fine-tuning-ready chat format

python
import json

def to_chat_format(record):
    """Convert a record to a simple {'messages': [...]} chat format
    commonly used for SFT / instruction-tuning pipelines."""
    return {
        "messages": [
            {"role": "user", "content": record["conversations"][0]["value"]},
            {"role": "assistant", "content": record["conversations"][1]["value"]},
        ]
    }

with open("merged_all_edu_prov_tax_misc_serial.jsonl", encoding="utf-8") as fin, \
     open("chat_format.jsonl", "w", encoding="utf-8") as fout:
    for line in fin:
        record = json.loads(line)
        chat_record = to_chat_format(record)
        fout.write(json.dumps(chat_record, ensure_ascii=False) + "\n")

11. Suggested Use Cases

  • —Fine-tuning / instruction-tuning Nepali-language LLMs on factual, statistics-grounded Q&A
  • —Evaluating a model's ability to read structured government statistical tables and answer precisely (numeric fidelity, unit handling, year/date grounding)
  • —Building retrieval-augmented generation (RAG) evaluation sets for Nepali government finance data
  • —Studying low-resource-language (Nepali/Devanagari) instruction data characteristics
  • —Benchmarking numeral formatting and Devanagari-numeral (०-९) generation in Nepali LLMs

12. Known Limitations / Caveats

  • —Every conversation is single-turn — there is no multi-turn dialogue or follow-up questioning in this file.
  • —Answers are short, single-fact statements; the dataset does not contain explanatory, analytical, or multi-fact reasoning answers.
  • —The Department of Roads subset (51 rows) has source_name = null and license = null — treat these as needing separate license verification before redistribution.
  • —The language/script tags for the Department of Roads subset say en/Latn, but the actual question/answer text is Devanagari Nepali — always check the literal conversations text rather than relying solely on the language field.
  • —The "Intermediate Consumption" source title references India, but its Q&A content is about Nepali economic sectors — this looks like an inherited title from the upstream source and does not reflect the actual content of the questions/answers.
  • —Numeric answers are presented as plain figures without consistent unit disambiguation in every row (e.g., some rows omit whether values are in millions, billions, or raw currency units) — cross-check the source field/original publication when unit precision matters.

13. Field Quick-Reference (cheat sheet)

I want to...Use this field
Get the question textconversations[0]["value"]
Get the answer textconversations[1]["value"]
Group by question-pattern/templateStrip trailing _NNNN from id
Filter by topic domainmetadata_json.generation_domain (after parsing)
Filter by original publicationsource or source_name
Filter by publisher/agencysource_repo
Filter by languagelanguage (ne / en)
Filter by licenselicense / license_tier
Check intended answering behaviormetadata_json.behavior / behavior_definition