CoolFace
Datasetpublic

sabin1234/Nepal_CRS_Company_FAQ_Contraception_Family_Planning_Nepali_QA_Dataset

Nepal CRS Company FAQ — Contraception & Family Planning Nepali Q&A Dataset 1. Overview This dataset is a Nepali-language (Devanagari script) collection of question–answer pairs in ShareGPT format, covering frequently asked questions about contraception and family planning methods — oral pills, emergency contraceptive pills, injectables (e.g. DMPA), implants, IUDs, and condoms. The content originates from Nepal CRS Company, a well-established Nepali… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Nepal_CRS_Company_FAQ_Contraception_Family_Planning_Nepali_QA_Dataset.

sourceHugging Faceapache-2.0updated 18d agoView on Hugging Face
0likes46downloads
Dataset Card

Nepal CRS Company FAQ — Contraception & Family Planning Nepali Q&A Dataset

1. Overview

This dataset is a Nepali-language (Devanagari script) collection of question–answer pairs in ShareGPT format, covering frequently asked questions about contraception and family planning methods — oral pills, emergency contraceptive pills, injectables (e.g. DMPA), implants, IUDs, and condoms. The content originates from Nepal CRS Company, a well-established Nepali social-marketing organization focused on family-planning and reproductive-health products and education.

Each record is a single-turn conversation where a human asks a factual, often myth-busting question in Nepali about a specific contraceptive method's safety, effectiveness, or side effects, and the assistant ("gpt") gives a clear, evidence-based answer in Nepali — typically opening with a direct "हो"/"होइन" (yes/no) followed by supporting explanation.

PropertyValue
File namenepal_crs_company_faq_nepali.jsonl
File formatJSON Lines (.jsonl) — one JSON object per line
Total records (rows)53
Conversation formatShareGPT (conversations list with from/value pairs)
Turns per conversation2 (1 human turn + 1 gpt turn) — 100% of records
DomainHealth — Contraception / family planning
Content languageNepali (Devanagari script) — 100% verified
Data originNepal CRS Company FAQ content (marked as "real" — not synthetically generated)
License (per-record)Apache-2.0 (permissive)
Duplicate IDs found0 (all 53 id values unique)
Duplicate questions found0
Malformed / unparseable lines0

2. File & Record Structure

This dataset uses the same flat schema style seen in the Global Healthcare Jobs FAQ dataset — provenance metadata at the top level, with fine-grained generation metadata nested in a stringified metadata_json field, and no numeric length/complexity constraint fields in the metadata.

KeyTypeDescription
idstringUnique record identifier, format crs_faq_0XX (e.g. crs_faq_001).
conversationsarrayShareGPT-style conversation turns. Always exactly 2 turns.
conversations[0].fromstringAlways "human" — the question turn.
conversations[0].valuestringThe Nepali-language question about a contraceptive method.
conversations[1].fromstringAlways "gpt" — the answer turn.
conversations[1].valuestringA Nepali-language, evidence-based, myth-clarifying answer.
source / source_repostringBoth "Nepal CRS Company".
source_namestringMachine-style source identifier (snake_case).
source_configstringConfig label.
source_splitstringData split label.
source_revisionstringVersion tag.
source_row_idstringRow-level traceability ID (identical to top-level id).
languagestringISO language code.
language_codestringISO 639-3 code.
scriptstringScript tag.
licensestringLicense of the content.
license_tierstringLicense permissiveness bucket.
task_typestringNLP task type this record represents.
generation_typestringHow the record was produced.
conditionstringGeneration condition tag.
urlstringSource URL reference (always empty in this file).
metadata_jsonstringStringified JSON object with fine-grained generation metadata (parsed in Section 3).

2.1 Top-level provenance fields (all constant across the dataset)

FieldObserved value%
sourceNepal CRS Company100%
source_namenepal_crs_company_faq100%
source_repoNepal CRS Company100%
source_configdefault100%
source_splittrain100%
source_revisionv1100%
languagene100%
language_codenpi100%
scriptDeva (Devanagari)100%
licenseApache-2.0100%
license_tierpermissive100%
task_typeinstruction-following100%
generation_typereal100%
conditionreal100%
url"" (empty)100%

This dataset is tagged `generation_type: real` / `condition: real`, and its schema (flat structure, v1 revision, sequential crs_faq_0XX IDs, no numeric length metadata) closely mirrors the Global Healthcare Jobs FAQ dataset reviewed previously — both appear to have been produced by the same translation/localization pipeline.

2.2 metadata_json sub-fields (parsed)

FieldDescriptionObserved value(s)
generation_domainTop-level subject domain.स्वास्थ्य ("Health") — 100%
generation_categorySub-topic category.गर्भनिरोधक सम्बन्धी प्रायः सोधिने प्रश्न ("Frequently asked questions about contraception") — 100%
generation_sub_domainSpecific FAQ theme.मौखिक चक्की, आपतकालीन चक्की, इन्जेक्टेबल, इम्प्लान्ट, आईयूडी र कन्डम ("Oral pills, emergency pills, injectables, implants, IUDs, and condoms") — 100%
behaviorExpected assistant behavior.जानकारीमूलक उत्तर दिने ("Provide an informative answer") — 100%
behavior_definitionFull definition of the expected behavior."मूल प्रश्नको अर्थ र मूल उत्तरको तथ्य कायम राख्दै शुद्ध नेपालीमा रूपान्तरण गरिएको उत्तर दिने।" ("Give an answer translated into pure Nepali while preserving the meaning of the original question and the facts of the original answer.") — 100%
question_typeFormat classification.तथ्यमा आधारित प्रश्न ("Fact-based question") — 100%
content_languageDeclared content language.नेपाली ("Nepali") — 100%, accurate
content_scriptDeclared content script.देवनागरी ("Devanagari") — 100%, accurate
english_content_allowedWhether English content is permitted.False — 100%

As with Global Healthcare Jobs, the behavior definition's explicit phrasing about "translating into pure Nepali while preserving the original question's meaning and the original answer's facts" indicates this dataset was produced by translating an existing English-language contraception-FAQ resource into Nepali (this content strongly resembles the globally used WHO/FHI 360-style "Family Planning: A Global Handbook for Providers" myth-busting Q&A format), rather than being freshly authored or templated from structured data.


3. Question Distribution

3.1 By domain/category/sub-domain

The dataset is single-domain and single-topic by design:

LevelValueCount%
generation_domainस्वास्थ्य (Health)53100%
generation_categoryContraception FAQ53100%
generation_sub_domainOral pills, emergency pills, injectables, implants, IUDs, condoms53100%
question_typeतथ्यमा आधारित प्रश्न (Fact-based question)53100%

All questions are open-ended and explanatory (no MCQ format) — most follow a classic myth-vs-fact structure ("Does X method cause Y?"), and the assistant is expected to directly confirm or refute the claim with supporting evidence.

3.2 Thematic sub-topic breakdown (derived from question content)

While the formal generation_sub_domain field already names the six contraceptive method categories the dataset covers, question-level content clustering below shows how many questions actually reference each method (a question may reference more than one theme, so counts overlap):

Contraceptive method / themeExample keyword(s) matchedQuestions matched% of dataset
Pregnancy / fertility effects (cross-cutting)गर्भ / बाँझोपन / सन्तान2750.9%
Oral pillsमौखिक गर्भनिरोधक चक्की / चक्की1732.1%
Implantsइम्प्लान्ट1120.8%
IUDआईयूडी1018.9%
Side effects / safetyसाइड इफेक्ट / असर / सुरक्षित / हानि917.0%
Emergency contraceptionआपतकालीन815.1%
Injectables (e.g. DMPA)इन्जेक्टेबल / इन्जेक्सन713.2%
Condomsकन्डम59.4%
Birth defectsजन्म दोष47.5%
Breastfeeding compatibilityस्तनपान23.8%
Effectivenessप्रभावकारी / असरदार11.9%

50 of 53 questions (94.3%) matched at least one of the above keyword groups, showing this dataset is very tightly and consistently focused on its stated topic. Pregnancy/fertility-effect concerns are the most cross-cutting theme (present in about half the questions, often as a component of a larger question about a specific method), while oral pills and implants are the two most frequently discussed individual method categories. Common myth patterns addressed include: "does this method cause abortion / birth defects / infertility / cancer?", "is this method safe for smokers / breastfeeding mothers / HIV-positive women / heavier women?", and "how quickly does fertility return after stopping this method?" — a pattern consistent with addressing widespread misconceptions about contraception.

3.3 Length characteristics

MetricMinMedianMeanMax
Question — character length287081.3191
Question — word count41011.225
Answer — character length85244286.3635
Answer — word count113440.189
Answer — sentence count (approx.)133.47

Answers are moderate-length, multi-sentence explanations (mean 286.3 characters / 40.1 words) — longer than the Global Healthcare Jobs dataset's answers (mean 173.5 characters) but shorter than the Harley Street Institute dataset's answers (mean 357.2 characters), sitting in a similar range to the Fight Vitiligo dataset (mean 305.2 characters). This fits the pattern of a clear, evidence-grounded myth-busting explanation rather than either a one-line fact or a long advisory essay.

3.4 Answer format

Answers typically open with a direct confirmation or denial ("हो।" / "होइन।" — "Yes." / "No.") immediately followed by 2–4 sentences of supporting evidence or explanation, closely matching a classic public-health myth-busting FAQ format. No lettered options or MCQ structure is present anywhere in the dataset.


4. Behavior Distribution

The dataset targets a single, uniform assistant behavior:

BehaviorBehavior definitionCount%
जानकारीमूलक उत्तर दिने ("Provide an informative answer")"Give an answer translated into pure Nepali while preserving the meaning of the original question and the facts of the original answer."53100%

4.1 Behavior-related generation notes

As with the Global Healthcare Jobs dataset, this dataset's metadata_json does not define numeric constraints such as maximum_response_sentences, minimum_reasoning_dimensions, minimum_complexity_score, question_length, or response_length. The only explicit behavioral constraint present is:

ConstraintValueMeaning
english_content_allowedFalseNo English content permitted anywhere in the record — output must be pure Nepali.

4.2 Task type / generation type

FieldValueCount%
task_typeinstruction-following53100%
generation_typereal53100%
conditionreal53100%

5. Provenance & Licensing Distribution

FieldValueCount%
sourceNepal CRS Company53100%
source_namenepal_crs_company_faq53100%
source_repoNepal CRS Company53100%
source_splittrain53100%
source_revisionv153100%
licenseApache-2.053100%
license_tierpermissive53100%
  • —Single source: all 53 records trace back to one content source — Nepal CRS Company, a Nepali social-marketing/reproductive-health organization.
  • —License: Apache-2.0, matching the license tier used in the Fight Vitiligo, Global Healthcare Jobs, Harley Street Institute, and Healthy Choice datasets.
  • —Traceability: each record's source_row_id matches its id directly (crs_faq_001–crs_faq_053), giving a simple, sequential trace back to source ordering.
  • —No URL field populated: every record's url field is an empty string.

6. Language & Script Notes

This dataset's language metadata is accurate and consistent with its actual content, matching the pattern seen across the other "real"-content Nepali FAQ datasets:

FieldDeclared valueActual content observed
languageneNepali ✅
language_codenpiNepali (ISO 639-3: npi) ✅
scriptDevaDevanagari ✅
metadata_json.content_languageनेपाली (Nepali)Nepali ✅
metadata_json.content_scriptदेवनागरी (Devanagari)Devanagari ✅
metadata_json.english_content_allowedFalseNo English text found ✅

Verified finding: A full script scan of all 53 records confirms 100% of both questions and answers are written entirely in Devanagari script — zero Latin-only text in either turn. Medical/pharmacological terms are transliterated phonetically into Devanagari rather than translated (e.g. "डीएमपीए" for "DMPA," "आईयूडी" for "IUD," "इम्प्लान्ट" for "implant," "एन्टिरेट्रोभाइरल थेरापी" for "antiretroviral therapy") — a natural localization approach for specialized clinical/pharmaceutical terminology, not a script violation.


7. Content Coverage Notes

  • —Single health-domain focus: every question concerns contraception and family-planning methods, addressing the six method categories named in the source category field: oral pills, emergency contraceptive pills, injectables, implants, IUDs, and condoms.
  • —Myth-busting / evidence-based framing: the dataset systematically addresses common public misconceptions about contraception — abortion-inducing effects, birth defects, infertility, cancer risk, ectopic pregnancy risk, weight gain, and safety for specific populations (smokers, breastfeeding mothers, HIV-positive women, women who have never given birth, heavier women).
  • —STI/condom-specific content: a subset of questions also address condom effectiveness against pregnancy and HIV/STI transmission, including questions about breakage/slippage and anal-sex use.
  • —No duplicate records or questions: all 53 id values and all 53 question texts are unique.
  • —No malformed JSON lines: every line parses successfully with a consistent flat schema.
  • —Sequential, gap-free IDs: crs_faq_001 through crs_faq_053, matching the reported row count exactly.

8. Suggested Use Cases

  • —Fine-tuning or evaluating Nepali-language LLMs for reproductive health / family-planning education and counseling applications.
  • —Domain-adaptation for Nepali-language sexual and reproductive health chatbots aimed at debunking common contraception myths.
  • —Benchmarking evidence-based, myth-correction answer generation in Nepali (direct yes/no + supporting rationale format).
  • —Supporting public health messaging and outreach tools in Nepal, where misinformation about contraceptive methods can be a barrier to family-planning uptake.
  • —As part of a combined Nepali health-FAQ corpus alongside Fight Vitiligo, Global Healthcare Jobs, Harley Street Institute, and Healthy Choice Aesthetic Hospital datasets, to study cross-domain consistency in "real"/translated Nepali health FAQ content.

9. Known Limitations

  • —Small dataset size: only 53 records — useful for narrow fine-tuning/evaluation on this specific FAQ topic, but too small alone for broad instruction tuning.
  • —Single domain, single behavior: entirely focused on contraception/family-planning FAQ with one behavior type (जानकारीमूलक उत्तर दिने) — should be combined with other datasets for general-purpose or multi-domain instruction tuning.
  • —No length/complexity metadata: this dataset provides no question_length, response_length, maximum_response_sentences, minimum_reasoning_dimensions, or minimum_complexity_score fields, limiting fine-grained filtering by answer length/complexity.
  • —No explicit source URLs: the url field is empty for every record, so individual FAQ entries cannot be traced back to a specific original source page.
  • —Sensitive health topic: content covers sexual and reproductive health in direct, clinical terms; deployers should ensure downstream use contexts (e.g. age-appropriate chatbots) are suitable for this subject matter, and should pair this dataset with appropriate safety/moderation layers when used in consumer-facing products.
  • —Medical currency: contraceptive guidance can evolve as new products, formulations, and clinical guidelines emerge; the factual accuracy of specific claims should be periodically reverified against current medical guidance (e.g. WHO Medical Eligibility Criteria for Contraceptive Use).