sabin1234/NEPSE_Dividend_FAQ
π NEPSE Dividend FAQ (Nepali) β Dataset README "How much dividend did it give this time?" β the question every stock market watcher asks, now turned into 2,000 question-answer pairs. A simple, clean, 100% Nepali-language financial FAQ dataset that builds question-answer (Q&A) pairs about the total dividend and bonus share dividend of companies listed on NEPSE. This file is a full dissection of that exact dataset β every field, every pattern, everything. π TL;DRβ¦ See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/NEPSE_Dividend_FAQ.
π NEPSE Dividend FAQ (Nepali) β Dataset README
"How much dividend did it give this time?" β the question every stock market watcher asks, now turned into 2,000 question-answer pairs.
A simple, clean, 100% Nepali-language financial FAQ dataset that builds question-answer (Q&A) pairs about the total dividend and bonus share dividend of companies listed on NEPSE. This file is a full dissection of that exact dataset β every field, every pattern, everything.
π TL;DR (At a Glance)
ποΈ Schema β What's Inside Each Record
Each line is an independent JSON object, and all of them share the same structure:
{
"id": "nepse_dividend_0001",
"conversations": [
{"from": "human", "value": "What was ACLBSL's total dividend for fiscal year 2077/078?"},
{"from": "gpt", "value": "ACLBSL's total dividend for fiscal year 2077/078 was 21%."}
],
"source": "NEPSE Dividend History",
"language": "ne", "language_code": "npi", "script": "Deva",
"license": "Apache-2.0", "license_tier": "permissive",
"task_type": "instruction-following",
"generation_type": "real", "condition": "real",
"metadata_json": "{...}"
}Field-by-field status
Interesting fact: almost all dataset-metadata fields (source, license, language, tasktype, generationtype...) are 100% constant β meaning this is a homogeneous dataset built from a single "batch" using a single recipe. All the variation lives in the actual Q&A content β the company, the fiscal year, and the dividend percentage.
Inside metadata_json (nested) β also all constant!
{
"generation_domain": "Financial Services",
"generation_category": "NEPSE Dividend History",
"generation_sub_domain": "Company Dividends",
"behavior": "Providing factual dividend information",
"behavior_definition": "Directly answering the question using dividend data for the specified company and fiscal year.",
"question_type": "Fact-based question",
"content_language": "Nepali",
"content_script": "Devanagari",
"english_content_allowed": false
}All 9 of these sub-fields are identical across 2,000 out of 2,000 lines β the behavior is a single, focused capability: "providing fact-based company dividend information." There's no multi-behavior mixing, and no opinion/analysis-style questions β pure facts only.
π¬ Conversation Structure
- Each record has exactly 2 turns:
human(question) βgpt(answer). Never 3-turn, 1-turn, or follow-up β this pattern holds for 100% of 2,000 records. - There is no system prompt.
- Length is also very tight/consistent:
Short, direct, single-fact Q&A β this is not a long conversation or reasoning-chain dataset. It's a quick-recall FAQ style.
β Question Pattern Distribution β Exactly 6 Templates
Once you strip out the company code and fiscal year, all 2,000 questions fit into one of these 6 templates (100% coverage, no outliers):
Template 1 ββββββββββββββββββββββ 23.4%
Template 2 βββββββββββββββββββββ 22.4%
Template 3 ββββββββββββββββββββ 20.1%
Template 4 βββββββββββ 11.6%
Template 5 βββββββββββ 11.6%
Template 6 βββββββββββ 11.1%Templates 1β3 = 3 phrasings asking about Total Dividend (1,317 records β 65.9%) Templates 4β6 = 3 phrasings asking about Bonus Share dividend (683 records β 34.1%)
For each "content" type (total/bonus), exactly 3 syntactic variations are provided β to give sentence-level diversity. But semantically, they're all just two kinds of questions.
β Answer Pattern Distribution β Exactly 2 Templates
The answers are even cleaner β there are only 2 templates total:
No hedging, no extra context, no "this is an estimate"-style disclaimers β a straight fact β answer, in a deterministic format. This design reveals the dataset's purpose: "given a specific company + specific year, the model should give a clean, single-line, fact-based answer."
π’ Company Coverage
- 211 distinct company codes are covered (e.g., EBL, NABIL, ADBL, GBIME...).
- Each company has an average of ~9.5 questions, but the distribution is quite skewed β some companies have data spanning many years, others have only one or two.
Top 15 Companies by Question Count
Records-per-Company Density (histogram)
Long-tail distribution β 20 companies have only 1 question each, while top-tier companies (large banks/finance companies like EBL, NABIL, GBIME) have up to 30 years of history covered. This reflects real-world data availability β older/larger companies have more historical records available, while newer/smaller ones have less.
π Fiscal Year Coverage
Fiscal years from 2062/063 to 2082/083 (nearly a 20-year range) are covered β but density is higher toward recent years:
2067β069 ββββββββββββββββ
2070β072 βββββββββββββββββββββββ
2073β075 βββββββββββββββββββββββββββββββ
2076β078 ββββββββββββββββββββββββββββββββββββββ
2079β081 ββββββββββββββββββββββββββββββββββββββββ
2082β083 βββββββFY 2081/082 has the most records (206), while the earliest years (062-067) have very few (just 3 total) β a natural recency bias, since older data is less digitized/available.
π° Dividend % Distribution
- Median: 13.5% | Average: ~19.7% (pulled upward by outliers)
- Most frequent single values: 10% (124 times), 15% (120 times), 0% (114 times), 20% (102 times)
- 392 distinct % values appear in total β many of them include decimals (e.g., 10.53%, 21.05%, 15.79%), which look like exact figures from actual bonus-share conversions, not rounded numbers.
π₯ The Highest Dividend (an interesting outlier!)
UNL appears to have declared a total dividend of 990% in FY 2071/072 and 860% in FY 2070/071.
This could be a data-entry error (e.g., 99.0% mistakenly entered as 990%), or it could genuinely reflect a bonus-heavy year for a micro-finance-type company back then. If you plan to use this dataset for training/production, it's a good idea to audit such extreme outliers.
π Bonus Share vs. Total Dividend
- Only 1 total-dividend record shows 0% β almost every company/year has some form of dividend.
- But 113 bonus-share records show 0% β meaning many companies didn't distribute bonus shares in certain years (cash dividend only), which matches real NEPSE practice.
π Duplicates & Data Quality
Overall, this dataset is very clean β almost production-ready for building FAQ pairs. Aside from the ~990% outlier mentioned above, no major errors were found.
π¨ Design Philosophy (the likely thinking behind this dataset)
- Single, focused behavior β it does one thing: given a company + year, return the % dividend. No multi-task mixing.
- Controlled syntactic diversity, fixed semantics β 6 question templates but the same underlying meaning (3 total + 3 bonus) β this seems aimed at making the model phrasing-robust (able to answer correctly no matter how the same information is asked).
- Real-world grounding β
generation_type: realβ not synthetic/fictional data, but drawn from actual NEPSE dividend history (hence the outliers/uneven distribution β the real world is messy!). - Deterministic, hedge-free answers β no words like "approximately" or "estimated." This establishes raw ground truth for factual-recall tasks.
- Nepali-only, Devanagari-only β
english_content_allowed: falseβ no Roman/English mixed in anywhere, full linguistic purity in Nepali is maintained.
π§ͺ Sample Records
{"id": "nepse_dividend_0001",
"conversations": [
{"from": "human", "value": "What was ACLBSL's total dividend for fiscal year 2077/078?"},
{"from": "gpt", "value": "ACLBSL's total dividend for fiscal year 2077/078 was 21%."}
]}
{"id": "nepse_dividend_0231",
"conversations": [
{"from": "human", "value": "What is NICL's bonus share dividend percentage for fiscal year 2074/075?"},
{"from": "gpt", "value": "NICL's bonus share dividend for fiscal year 2074/075 was 10.53%."}
]}β οΈ Things to Watch Out For When Using This Dataset
- Coverage is uneven β some companies/years have a lot of data, others just one record. Watch out for class imbalance when fine-tuning.
- Only company codes are given, not full names (e.g., "EBL" β presumably Everest Bank Limited) β if full names are needed, a separate mapping table is required.
- Outliers like the 990% figure should be audited.
- 2-turn, single-fact QA only β this dataset doesn't teach multi-turn conversation, reasoning, or comparative analysis ("which company paid more?"). If that kind of capability is needed, it should be combined with another dataset.
π License
Apache-2.0 (permissive) β free for both commercial and non-commercial use, with attribution.
This README was auto-generated β based on a programmatic analysis of the full 2,000 records (all figures calculated from the actual dataset, not estimated).
