CoolFace
Datasetpublic

sabin1234/Healthy_Choice_Aesthetic_Hospital_FAQ

Healthy Choice Aesthetic Hospital FAQ — Pure Nepali Q&A Dataset 1. Overview This dataset is a pure Nepali-language (Devanagari script) collection of question–answer pairs in ShareGPT format, covering frequently asked questions about Healthy Choice Aesthetic Hospital — a real aesthetic/cosmetic and general specialty hospital located in Kathmandu, Nepal. Topics span hospital directions/location, hair transplant, skin treatments, plastic surgery, laser procedures… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Healthy_Choice_Aesthetic_Hospital_FAQ.

sourceHugging Faceapache-2.0updated 17d agoView on Hugging Face
0likes40downloads
Dataset Card

Healthy Choice Aesthetic Hospital FAQ — Pure Nepali Q&A Dataset

1. Overview

This dataset is a pure Nepali-language (Devanagari script) collection of question–answer pairs in ShareGPT format, covering frequently asked questions about Healthy Choice Aesthetic Hospital — a real aesthetic/cosmetic and general specialty hospital located in Kathmandu, Nepal. Topics span hospital directions/location, hair transplant, skin treatments, plastic surgery, laser procedures, dental, cardiology, and other clinical services offered by the hospital.

Each record is a single-turn conversation where a human asks a practical, patient-facing question in Nepali, and the assistant ("gpt") gives a concise, fact-based answer in Nepali.

PropertyValue
File namehealthy_choice_faq_suddha_nepali.jsonl
File formatJSON Lines (.jsonl) — one JSON object per line
Total records (rows)101
Conversation formatShareGPT (conversations list with from/value pairs)
Turns per conversation2 (1 human turn + 1 gpt turn) — 100% of records
DomainHealth — Aesthetic/cosmetic hospital FAQ (Healthy Choice Aesthetic Hospital, Kathmandu)
Content languageNepali (Devanagari script) — 100% verified
Underlying data sourceCohere Labs' Aya Dataset (CohereLabs/aya_dataset), Nepali subset
License (per-record)Apache-2.0 (permissive)
Duplicate IDs found0 (all 101 id values unique)
Duplicate questions found0
Malformed / unparseable lines0

2. File & Record Structure

This dataset uses the same flat schema style seen in the other three "sourced/real-content" FAQ datasets reviewed previously (Fight Vitiligo, Global Healthcare Jobs, Harley Street Institute), with provenance metadata at the top level and fine-grained generation metadata nested in a stringified metadata_json field.

KeyTypeDescription
idstringUnique record identifier, ShareGPT-style hash format sg_<32-char hex> (e.g. sg_900a97b9e2420ec424d75ca45ea6aa76).
conversationsarrayShareGPT-style conversation turns. Always exactly 2 turns.
conversations[0].fromstringAlways "human" — the question turn.
conversations[0].valuestringThe Nepali-language question about the hospital, its location, or its services.
conversations[1].fromstringAlways "gpt" — the answer turn.
conversations[1].valuestringA Nepali-language factual answer (length varies widely — see Section 3.3).
sourcestringOrigin dataset path.
source_namestringHuman-readable source name.
source_repostringRepository the source data was pulled from.
source_configstringConfig label.
source_splitstringData split label.
source_revisionstringFixed revision hash shared by all records in this file.
source_row_idstringRow-level traceability ID, format <id>:1.
languagestringISO language code.
language_codestringISO 639-3 code.
scriptstringScript tag.
licensestringLicense of the content.
license_tierstringLicense permissiveness bucket.
task_typestringNLP task type this record represents.
generation_typestringHow the record was produced.
conditionstringGeneration condition tag.
urlstringSource URL reference (always empty in this file).
metadata_jsonstringStringified JSON object with fine-grained generation metadata (parsed in Section 3).

2.1 Top-level provenance fields (all constant across the dataset)

FieldObserved value%
sourceCohereLabs/aya_dataset:default:train100%
source_nameaya_human_nepali100%
source_repoCohereLabs/aya_dataset100%
source_configdefault100%
source_splittrain100%
source_revisionf9ea04583f02a8f86404ff6c58bf75fe637df8a2100%
languagene100%
language_codenpi100%
scriptDeva (Devanagari)100%
licenseApache-2.0100%
license_tierpermissive100%
task_typeinstruction-following100%
generation_typesynthetic100%
conditionsynthetic100%
url"" (empty)100%

Notable difference from the other three "sourced FAQ" datasets reviewed: this is the only one whose underlying source traces to a well-known public multilingual instruction dataset — Cohere Labs' Aya Dataset (aya_human_nepali subset), rather than a company/institute's own FAQ page. It is also the only one of the four tagged generation_type: synthetic / condition: synthetic — matching the pattern seen in the earlier BFI banking dataset rather than the "real"/"original" tags seen in Fight Vitiligo, Global Healthcare Jobs, and Harley Street Institute.

2.2 metadata_json sub-fields (parsed)

FieldDescriptionObserved value(s)
generation_domainTop-level subject domain.स्वास्थ्य ("Health") — 100%
generation_categorySub-topic category.एस्थेटिक अस्पताल प्रश्नोत्तर ("Aesthetic hospital Q&A") — 100%
generation_sub_domainSpecific entity/theme.हेल्दी च्वाइस एस्थेटिक अस्पताल ("Healthy Choice Aesthetic Hospital") — 100%
behaviorExpected assistant behavior.छोटो उत्तर माग्ने ("Request a short answer") — 100%
behavior_definitionFull definition of the expected behavior."मुख्य जानकारी कायम राख्दै संक्षिप्त उत्तर दिने।" ("Give a concise answer while retaining the key information.") — 100%
question_typeFormat classification.प्रश्नोत्तर प्रश्न ("Question-answer question") — 100%
question_lengthDeclared character-length bucket for the question.५० देखि ३०० अक्षर ("50 to 300 characters") — 100%
response_lengthDeclared character-length bucket for the answer.५० देखि ८०० अक्षर ("50 to 800 characters") — 100%
maximum_response_sentencesMax sentences allowed in the answer.15 — 100%
minimum_reasoning_dimensionsMinimum reasoning dimensions required.1 — 100%
minimum_complexity_scoreMinimum complexity score assigned.3 — 100%
content_languageDeclared content language.नेपाली ("Nepali") — 100%, accurate
content_scriptDeclared content script.देवनागरी ("Devanagari") — 100%, accurate
english_content_allowedWhether English content is permitted.False — 100%

This dataset's generation constraints allow a notably wider and larger permitted answer length (up to 800 characters / 15 sentences) than the other three "real"/"original" FAQ datasets, and the actual maximum observed answer (1,223 characters — see Section 3.3) exceeds even that declared 800-character upper bound, indicating the length bucket is a soft/approximate target rather than a hard limit.


3. Question Distribution

3.1 By domain/category/sub-domain

The dataset is single-domain, single-entity, and single-topic by design — every record belongs to the same taxonomy path:

LevelValueCount%
generation_domainस्वास्थ्य (Health)101100%
generation_categoryएस्थेटिक अस्पताल प्रश्नोत्तर (Aesthetic hospital Q&A)101100%
generation_sub_domainहेल्दी च्वाइस एस्थेटिक अस्पताल (Healthy Choice Aesthetic Hospital)101100%
question_typeप्रश्नोत्तर प्रश्न (Question-answer question)101100%

Unlike the earlier BFI banking MCQ dataset, there is no multiple-choice format here — every answer is free-text, and unlike Fight Vitiligo/Global Healthcare Jobs/Harley Street Institute, all questions are anchored to one single named business (a specific hospital) rather than a general topic area.

3.2 Thematic sub-topic breakdown (derived from question content)

The table below is a keyword-based approximate clustering of the actual question text (a question may match more than one theme, so counts can overlap):

Sub-theme (approx.)Example keyword(s) matchedQuestions matched% of dataset
Surgery / proceduresशल्यक्रिया / अपरेसन2221.8%
Hair transplantकपाल प्रत्यारोपण / कपाल1211.9%
Opening hours / timingसमय / खुल्छ / बन्द76.9%
Treatment results / side effectsनतिजा / साइड इफेक्ट / असर76.9%
Diet / nutritionआहार / पोषण65.9%
Dentalदन्त55.0%
Cost / priceखर्च / शुल्क / मूल्य / लागत44.0%
Skin treatmentछाला44.0%
Department / services offeredविभाग / सेवा33.0%
Doctor / specialistडाक्टर / विशेषज्ञ / चिकित्सक22.0%
Location / directionsकहाँ अवस्थित / नक्सा / बाटो / निर्देशन11.0%
Appointment / bookingअपोइन्टमेन्ट / बुकिङ / भेटघाट11.0%
Other / unmatched by keyword list—4544.6%

56 of 101 questions (55.4%) matched at least one of the above keyword groups. Surgery/procedures and hair transplant are the dominant identifiable themes — reflecting the hospital's core aesthetic-medicine service lines (hair transplant methods like FUE/DHI, laser tattoo removal, non-surgical fat reduction/cryolipolysis, permanent hair removal, mole removal, acne and melasma treatment). Many "other/unmatched" questions cover further procedure-specific detail not captured by the simple keyword list (e.g. HydraFacial benefits, andrology consultations, donor/recipient area terminology in hair transplant, helmet-wearing safety timing after transplant).

3.3 Length characteristics

MetricMinMedianMeanMax
Question — character length184646.792
Question — word count377.514
Answer — character length43131181.01,223
Answer — word count62026.6185
Answer — sentence count (approx.)111.918

Answer lengths are highly skewed — most answers are short (median 131 characters, 20 words), but a small number of records are much longer, dominated by one clear outlier: the hospital's directions/location answer (1,223 characters, turn-by-turn navigation instructions from two different Kathmandu landmarks — Naxal/Bhatbhateni and Chabahil). Excluding that single outlier, the next-longest answers (669, 623, 558, 514 characters) relate to detailed procedural explanations (e.g. andrologist consultation timing, hair transplant methods comparison). This wide variance (max 1,223 vs. declared bucket ceiling of 800 characters) shows the dataset mixes very short factual answers with occasional long, structured, multi-part explanations as needed by the question.

3.4 Answer format

Answers are free-form and practical, ranging from single-sentence direct facts (e.g. cost or session-count answers) to long, structured, multi-step explanations (e.g. the detailed floor-by-floor, landmark-by-landmark directions to the hospital, which reads almost like a step-by-step wayfinding guide). No lettered options or MCQ structure is present anywhere in the dataset.


4. Behavior Distribution

The dataset targets a single, uniform assistant behavior:

BehaviorBehavior definitionCount%
छोटो उत्तर माग्ने ("Request a short answer")"Give a concise answer while retaining the key information."101100%

4.1 Behavior-related generation constraints

ConstraintValueMeaning
maximum_response_sentences15A relatively generous sentence cap — allows multi-step or multi-part answers (e.g. the long directions answer) while still bounding total length.
minimum_reasoning_dimensions1Each answer requires only 1 reasoning dimension — direct factual lookup/recall, not multi-fact combination.
minimum_complexity_score3A mid-low complexity tier — higher than the Fight Vitiligo dataset's score of 1, but well below the BFI banking dataset's score of 5.
question_length50–300 characters (declared)Actual observed range: 18–92 characters — narrower than the declared bucket.
response_length50–800 characters (declared)Actual observed range: 43–1,223 characters — exceeds the declared upper bound for the longest (directions) answer.
english_content_allowedFalseNo English content permitted anywhere in the record — output must be pure Nepali.

4.2 Task type / generation type

FieldValueCount%
task_typeinstruction-following101100%
generation_typesynthetic101100%
conditionsynthetic101100%

Despite being labeled synthetic (as opposed to real/original in the other business-FAQ-style datasets reviewed), the content is clearly grounded in real, specific facts about an actual hospital (its physical layout, block/floor structure, nearby landmarks, and service offerings), suggesting the Q&A pairs were synthetically generated from real underlying hospital information rather than being freely invented.


5. Provenance & Licensing Distribution

FieldValueCount%
sourceCohereLabs/aya_dataset:default:train101100%
source_nameaya_human_nepali101100%
source_repoCohereLabs/aya_dataset101100%
source_splittrain101100%
source_revisionf9ea04583f02a8f86404ff6c58bf75fe637df8a2101100%
licenseApache-2.0101100%
license_tierpermissive101100%
  • —Single source: all 101 records trace back to the same underlying source snapshot — the Nepali-language human-written subset of Cohere Labs' publicly available Aya Dataset, a large multilingual instruction-tuning dataset.
  • —License: Apache-2.0, matching the license tier used in the Fight Vitiligo, Global Healthcare Jobs, and Harley Street Institute datasets.
  • —Traceability: source_row_id mirrors the top-level id with a :1 suffix, and all records share one fixed source_revision hash — indicating this file is a single filtered/extracted batch (the Healthy Choice Aesthetic Hospital-related subset) pulled from the larger Aya Dataset Nepali collection.
  • —No URL field populated: every record's url field is an empty string.

6. Language & Script Notes

This dataset's language metadata is accurate and consistent with its actual content, matching the pattern seen in Fight Vitiligo, Global Healthcare Jobs, and Harley Street Institute:

FieldDeclared valueActual content observed
languageneNepali ✅
language_codenpiNepali (ISO 639-3: npi) ✅
scriptDevaDevanagari ✅
metadata_json.content_languageनेपाली (Nepali)Nepali ✅
metadata_json.content_scriptदेवनागरी (Devanagari)Devanagari ✅
metadata_json.english_content_allowedFalseNo English text found ✅

Verified finding: A full script scan of all 101 records confirms 100% of both questions and answers are written entirely in Devanagari script — zero Latin-only text in either turn. This matches the file name (suddha_nepali, i.e. "pure Nepali") accurately. As in the other real-content datasets, some medical/procedural terminology is transliterated phonetically into Devanagari rather than translated (e.g. "क्रायोलिपोलिसिस" for "cryolipolysis," "हाइड्राफेशियल" for "HydraFacial," "एन्ड्रोलोजिस्ट" for "andrologist," "एफयुई," "डीएचआई" for hair-transplant technique acronyms FUE/DHI) — a natural localization choice for specialized clinical terms, not a script violation.


7. Content Coverage Notes

  • —Single-institution focus: every question concerns one specific, named hospital — Healthy Choice Aesthetic Hospital in Kathmandu — including its physical location/directions, floor-by-floor department layout, and service offerings.
  • —Service lines covered: hair transplant (FUE, DHI, Instant DHT methods), skin treatments (melasma, acne, moles, HydraFacial), laser procedures (tattoo removal, permanent hair removal via Soprano), non-surgical body-fat reduction (cryolipolysis), plastic surgery, dental, cardiology, ophthalmology, and general andrology/consultation guidance.
  • —Detailed real-world grounding: the directions/location answer includes specific, verifiable local landmarks (Bhatbhateni supermarket, Nabil Bank, Chabahil, Om Hospital, "yellow bridge") and block/floor department mapping — indicating this content is grounded in the hospital's actual physical layout rather than being generic template text.
  • —No duplicate records or questions: all 101 id values and all 101 question texts are unique.
  • —No malformed JSON lines: every line parses successfully with a consistent flat schema.

8. Suggested Use Cases

  • —Fine-tuning or evaluating Nepali-language LLMs for local business/healthcare-facility customer-support chatbots (location/directions, service listings, pricing FAQs).
  • —Domain-adaptation for Nepali-language hospital or clinic virtual assistants, particularly for aesthetic/cosmetic and hair-restoration services.
  • —Benchmarking variable-length factual answer generation in Nepali — from single-fact short answers to long, structured, step-by-step directions.
  • —As a template for building similarly structured single-business FAQ datasets for other Nepali healthcare facilities, using the same Aya-Dataset-derived schema.
  • —As part of a combined Nepali health-FAQ corpus alongside Fight Vitiligo, Global Healthcare Jobs, and Harley Street Institute datasets, to study variation in answer style, length distribution, and generation-type labeling (synthetic vs. real/original) across differently-sourced Nepali health content.

9. Known Limitations

  • —Small dataset size: only 101 records — useful for narrow fine-tuning/evaluation on this specific hospital's FAQ, but too small alone for broad instruction tuning.
  • —Single domain, single entity, single behavior: entirely focused on one named hospital's FAQ with one behavior type (छोटो उत्तर माग्ने) — should be combined with other datasets for general-purpose or multi-domain instruction tuning.
  • —Highly skewed answer length: a small number of long answers (up to 1,223 characters) coexist with many very short ones (as few as 43 characters); models trained naively on this distribution may need length-aware sampling or weighting to avoid over-indexing on the outlier long-form directions answer.
  • —Declared length bucket exceeded: the metadata's stated response_length ceiling (800 characters) does not hold for at least one record (1,223 characters), so this field should be treated as an approximate target rather than a guaranteed hard bound.
  • —Highly location/time-specific content: directions, landmarks, and possibly pricing information reflect a specific point in time and physical layout; such details can become outdated if the hospital relocates, renovates, or changes pricing.
  • —No explicit source URLs: the url field is empty for every record, so individual FAQ entries cannot be traced back to a specific original web page.