sabin1234/Healthy_Choice_Aesthetic_Hospital_FAQ
Healthy Choice Aesthetic Hospital FAQ — Pure Nepali Q&A Dataset 1. Overview This dataset is a pure Nepali-language (Devanagari script) collection of question–answer pairs in ShareGPT format, covering frequently asked questions about Healthy Choice Aesthetic Hospital — a real aesthetic/cosmetic and general specialty hospital located in Kathmandu, Nepal. Topics span hospital directions/location, hair transplant, skin treatments, plastic surgery, laser procedures… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Healthy_Choice_Aesthetic_Hospital_FAQ.
Healthy Choice Aesthetic Hospital FAQ — Pure Nepali Q&A Dataset
1. Overview
This dataset is a pure Nepali-language (Devanagari script) collection of question–answer pairs in ShareGPT format, covering frequently asked questions about Healthy Choice Aesthetic Hospital — a real aesthetic/cosmetic and general specialty hospital located in Kathmandu, Nepal. Topics span hospital directions/location, hair transplant, skin treatments, plastic surgery, laser procedures, dental, cardiology, and other clinical services offered by the hospital.
Each record is a single-turn conversation where a human asks a practical, patient-facing question in Nepali, and the assistant ("gpt") gives a concise, fact-based answer in Nepali.
2. File & Record Structure
This dataset uses the same flat schema style seen in the other three "sourced/real-content" FAQ datasets reviewed previously (Fight Vitiligo, Global Healthcare Jobs, Harley Street Institute), with provenance metadata at the top level and fine-grained generation metadata nested in a stringified metadata_json field.
2.1 Top-level provenance fields (all constant across the dataset)
Notable difference from the other three "sourced FAQ" datasets reviewed: this is the only one whose underlying source traces to a well-known public multilingual instruction dataset — Cohere Labs' Aya Dataset (aya_human_nepali subset), rather than a company/institute's own FAQ page. It is also the only one of the four tagged generation_type: synthetic / condition: synthetic — matching the pattern seen in the earlier BFI banking dataset rather than the "real"/"original" tags seen in Fight Vitiligo, Global Healthcare Jobs, and Harley Street Institute.
2.2 metadata_json sub-fields (parsed)
This dataset's generation constraints allow a notably wider and larger permitted answer length (up to 800 characters / 15 sentences) than the other three "real"/"original" FAQ datasets, and the actual maximum observed answer (1,223 characters — see Section 3.3) exceeds even that declared 800-character upper bound, indicating the length bucket is a soft/approximate target rather than a hard limit.
3. Question Distribution
3.1 By domain/category/sub-domain
The dataset is single-domain, single-entity, and single-topic by design — every record belongs to the same taxonomy path:
Unlike the earlier BFI banking MCQ dataset, there is no multiple-choice format here — every answer is free-text, and unlike Fight Vitiligo/Global Healthcare Jobs/Harley Street Institute, all questions are anchored to one single named business (a specific hospital) rather than a general topic area.
3.2 Thematic sub-topic breakdown (derived from question content)
The table below is a keyword-based approximate clustering of the actual question text (a question may match more than one theme, so counts can overlap):
56 of 101 questions (55.4%) matched at least one of the above keyword groups. Surgery/procedures and hair transplant are the dominant identifiable themes — reflecting the hospital's core aesthetic-medicine service lines (hair transplant methods like FUE/DHI, laser tattoo removal, non-surgical fat reduction/cryolipolysis, permanent hair removal, mole removal, acne and melasma treatment). Many "other/unmatched" questions cover further procedure-specific detail not captured by the simple keyword list (e.g. HydraFacial benefits, andrology consultations, donor/recipient area terminology in hair transplant, helmet-wearing safety timing after transplant).
3.3 Length characteristics
Answer lengths are highly skewed — most answers are short (median 131 characters, 20 words), but a small number of records are much longer, dominated by one clear outlier: the hospital's directions/location answer (1,223 characters, turn-by-turn navigation instructions from two different Kathmandu landmarks — Naxal/Bhatbhateni and Chabahil). Excluding that single outlier, the next-longest answers (669, 623, 558, 514 characters) relate to detailed procedural explanations (e.g. andrologist consultation timing, hair transplant methods comparison). This wide variance (max 1,223 vs. declared bucket ceiling of 800 characters) shows the dataset mixes very short factual answers with occasional long, structured, multi-part explanations as needed by the question.
3.4 Answer format
Answers are free-form and practical, ranging from single-sentence direct facts (e.g. cost or session-count answers) to long, structured, multi-step explanations (e.g. the detailed floor-by-floor, landmark-by-landmark directions to the hospital, which reads almost like a step-by-step wayfinding guide). No lettered options or MCQ structure is present anywhere in the dataset.
4. Behavior Distribution
The dataset targets a single, uniform assistant behavior:
4.1 Behavior-related generation constraints
4.2 Task type / generation type
Despite being labeled synthetic (as opposed to real/original in the other business-FAQ-style datasets reviewed), the content is clearly grounded in real, specific facts about an actual hospital (its physical layout, block/floor structure, nearby landmarks, and service offerings), suggesting the Q&A pairs were synthetically generated from real underlying hospital information rather than being freely invented.
5. Provenance & Licensing Distribution
- Single source: all 101 records trace back to the same underlying source snapshot — the Nepali-language human-written subset of Cohere Labs' publicly available Aya Dataset, a large multilingual instruction-tuning dataset.
- License: Apache-2.0, matching the license tier used in the Fight Vitiligo, Global Healthcare Jobs, and Harley Street Institute datasets.
- Traceability:
source_row_idmirrors the top-levelidwith a:1suffix, and all records share one fixedsource_revisionhash — indicating this file is a single filtered/extracted batch (the Healthy Choice Aesthetic Hospital-related subset) pulled from the larger Aya Dataset Nepali collection. - No URL field populated: every record's
urlfield is an empty string.
6. Language & Script Notes
This dataset's language metadata is accurate and consistent with its actual content, matching the pattern seen in Fight Vitiligo, Global Healthcare Jobs, and Harley Street Institute:
Verified finding: A full script scan of all 101 records confirms 100% of both questions and answers are written entirely in Devanagari script — zero Latin-only text in either turn. This matches the file name (suddha_nepali, i.e. "pure Nepali") accurately. As in the other real-content datasets, some medical/procedural terminology is transliterated phonetically into Devanagari rather than translated (e.g. "क्रायोलिपोलिसिस" for "cryolipolysis," "हाइड्राफेशियल" for "HydraFacial," "एन्ड्रोलोजिस्ट" for "andrologist," "एफयुई," "डीएचआई" for hair-transplant technique acronyms FUE/DHI) — a natural localization choice for specialized clinical terms, not a script violation.
7. Content Coverage Notes
- Single-institution focus: every question concerns one specific, named hospital — Healthy Choice Aesthetic Hospital in Kathmandu — including its physical location/directions, floor-by-floor department layout, and service offerings.
- Service lines covered: hair transplant (FUE, DHI, Instant DHT methods), skin treatments (melasma, acne, moles, HydraFacial), laser procedures (tattoo removal, permanent hair removal via Soprano), non-surgical body-fat reduction (cryolipolysis), plastic surgery, dental, cardiology, ophthalmology, and general andrology/consultation guidance.
- Detailed real-world grounding: the directions/location answer includes specific, verifiable local landmarks (Bhatbhateni supermarket, Nabil Bank, Chabahil, Om Hospital, "yellow bridge") and block/floor department mapping — indicating this content is grounded in the hospital's actual physical layout rather than being generic template text.
- No duplicate records or questions: all 101
idvalues and all 101 question texts are unique. - No malformed JSON lines: every line parses successfully with a consistent flat schema.
8. Suggested Use Cases
- Fine-tuning or evaluating Nepali-language LLMs for local business/healthcare-facility customer-support chatbots (location/directions, service listings, pricing FAQs).
- Domain-adaptation for Nepali-language hospital or clinic virtual assistants, particularly for aesthetic/cosmetic and hair-restoration services.
- Benchmarking variable-length factual answer generation in Nepali — from single-fact short answers to long, structured, step-by-step directions.
- As a template for building similarly structured single-business FAQ datasets for other Nepali healthcare facilities, using the same Aya-Dataset-derived schema.
- As part of a combined Nepali health-FAQ corpus alongside Fight Vitiligo, Global Healthcare Jobs, and Harley Street Institute datasets, to study variation in answer style, length distribution, and generation-type labeling (
syntheticvs.real/original) across differently-sourced Nepali health content.
9. Known Limitations
- Small dataset size: only 101 records — useful for narrow fine-tuning/evaluation on this specific hospital's FAQ, but too small alone for broad instruction tuning.
- Single domain, single entity, single behavior: entirely focused on one named hospital's FAQ with one behavior type (
छोटो उत्तर माग्ने) — should be combined with other datasets for general-purpose or multi-domain instruction tuning. - Highly skewed answer length: a small number of long answers (up to 1,223 characters) coexist with many very short ones (as few as 43 characters); models trained naively on this distribution may need length-aware sampling or weighting to avoid over-indexing on the outlier long-form directions answer.
- Declared length bucket exceeded: the metadata's stated
response_lengthceiling (800 characters) does not hold for at least one record (1,223 characters), so this field should be treated as an approximate target rather than a guaranteed hard bound. - Highly location/time-specific content: directions, landmarks, and possibly pricing information reflect a specific point in time and physical layout; such details can become outdated if the hospital relocates, renovates, or changes pricing.
- No explicit source URLs: the
urlfield is empty for every record, so individual FAQ entries cannot be traced back to a specific original web page.
