sabin1234/Global_Healthcare_Jobs_FAQ
Global Healthcare Jobs FAQ — UK NHS Careers Nepali Q&A Dataset 1. Overview This dataset is a Nepali-language (Devanagari script) collection of question–answer pairs in ShareGPT format, covering frequently asked questions about working in the UK's National Health Service (NHS) — aimed specifically at international/foreign healthcare workers (nurses, doctors, dentists, and other clinicians) seeking employment, registration, visas, and career guidance in the United… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Global_Healthcare_Jobs_FAQ.
Global Healthcare Jobs FAQ — UK NHS Careers Nepali Q&A Dataset
1. Overview
This dataset is a Nepali-language (Devanagari script) collection of question–answer pairs in ShareGPT format, covering frequently asked questions about working in the UK's National Health Service (NHS) — aimed specifically at international/foreign healthcare workers (nurses, doctors, dentists, and other clinicians) seeking employment, registration, visas, and career guidance in the United Kingdom's healthcare sector.
Each record is a single-turn conversation where a human asks a practical, career-related question in Nepali about UK NHS job-seeking, and the assistant ("gpt") gives a short, factual, informative answer in Nepali.
2. File & Record Structure
This dataset uses the same flat schema style seen in the Fight Vitiligo FAQ dataset — most provenance metadata sits at the top level, with fine-grained generation metadata nested in a stringified metadata_json field.
2.1 Top-level provenance fields (all constant across the dataset)
As with the Fight Vitiligo dataset, this dataset is tagged `generation_type: real` / `condition: real`, indicating it is derived from genuine FAQ content rather than being purely template-generated.
2.2 metadata_json sub-fields (parsed)
Notable schema difference: unlike the Fight Vitiligo dataset, this dataset's metadata_json does not include question_length, response_length, maximum_response_sentences, minimum_reasoning_dimensions, or minimum_complexity_score fields — those length/complexity constraints are simply absent here, suggesting this dataset was generated/curated under a lighter-weight metadata schema (likely an earlier or simpler pipeline version, or a straight translation task — note the behavior definition explicitly says the answer is "translated into pure Nepali... preserving the original question's meaning and the original answer's facts," implying the source FAQ content was originally in English and has been translated to Nepali, unlike the Vitiligo dataset's more generic "give a fact-based answer" framing).
3. Question Distribution
3.1 By domain/category/sub-domain
The dataset is single-domain and single-topic by design — every record belongs to the same taxonomy path:
All questions are open-ended and explanatory (no MCQ format) — the assistant is expected to provide a direct, factual, career-advice-style answer.
3.2 Thematic sub-topic breakdown (derived from question content)
The formal taxonomy is uniform, but the actual question content clusters around several recognizable career-related sub-themes. The table below is a keyword-based approximate clustering (a question may match more than one theme, so counts can overlap):
52 of 97 questions (53.6%) matched at least one of the above keyword groups; the remainder cover related but more loosely-worded aspects of the UK healthcare job-search journey (e.g. career progression, professional networking, specific professional-council processes like changing a nursing PIN, sponsor licences, medical specialty registration, competitiveness of roles). Registration/licensing and qualification/experience recognition are the most prominent explicit themes, reflecting the central pain point for foreign-trained clinicians: getting credentials recognized by UK regulatory bodies (e.g. the Nursing and Midwifery Council for nurses, the General Medical Council for doctors) before they can legally practice.
3.3 Length characteristics
Answers here are noticeably shorter than the Fight Vitiligo dataset's (mean 173.5 characters / 23.4 words here, vs. mean 305.2 characters / 45.6 words for Vitiligo) — typically 1–2 concise sentences giving a direct, practical answer rather than a multi-sentence explanatory paragraph. This fits a "quick FAQ / career-tips" style rather than in-depth medical education content.
3.4 Answer format
Answers are free-form, direct, and practical — short factual statements or brief pieces of guidance (e.g. naming example recruitment agencies, confirming a process exists, or giving a concise yes/no-plus-explanation). No lettered options or MCQ structure is present anywhere in the dataset.
4. Behavior Distribution
The dataset targets a single, uniform assistant behavior:
4.1 Behavior-related generation notes
Unlike the two previously reviewed datasets, this dataset's metadata_json does not define numeric constraints such as maximum_response_sentences, minimum_reasoning_dimensions, or minimum_complexity_score — these fields are simply not present in any record. The only explicit behavioral constraint is:
The behavior definition's phrasing ("preserving the meaning of the original question and the facts of the original answer... translated into pure Nepali") strongly suggests this dataset was produced by translating an existing English-language NHS-careers FAQ resource into Nepali, rather than being generated from scratch or templated from structured data (contrast with the BFI banking dataset, which was templated from a structured registry, or the Vitiligo dataset, which was framed as direct fact-based Q&A generation).
4.2 Task type / generation type
5. Provenance & Licensing Distribution
- Single source: all 97 records trace back to one content source ("Global Healthcare Jobs" — a UK-focused healthcare recruitment/career-advice publisher).
- License: Apache-2.0, same permissive license tier as the Fight Vitiligo dataset (differs from the CC0-1.0 used in the BFI banking dataset). Apache-2.0 permits reuse and modification but expects attribution/notice preservation.
- Full traceability: each record's
source_row_idmatches itsiddirectly (global_healthcare_jobs_001–global_healthcare_jobs_097), giving a simple, sequential trace back to source ordering. - No URL field populated: every record's
urlfield is an empty string, so there is no direct link back to a specific original FAQ page per record.
6. Language & Script Notes
This dataset's language metadata is accurate and consistent with its actual content, matching the pattern seen in the Fight Vitiligo dataset:
Verified finding: A full script scan of all 97 records confirms 100% of both questions and answers are written entirely in Devanagari script — zero Latin-only text in either turn. Note, however, that several UK-specific institutional terms are transliterated into Nepali phonetically rather than translated (e.g. "राष्ट्रिय स्वास्थ्य सेवा" for "National Health Service," "नर्सिङ तथा मिडवाइफरी परिषद" for "Nursing and Midwifery Council," "सामान्य चिकित्सा परिषद" for "General Medical Council") — this is a translation/localization choice, not a script violation, since the surface text remains fully Devanagari.
7. Content Coverage Notes
- Single career-domain focus: every question concerns working in the UK NHS as a foreign/international healthcare professional — spanning nurses, doctors, dentists, and other clinical roles.
- Regulatory bodies referenced: the Nursing and Midwifery Council (NMC) for nurses, and the General Medical Council (GMC) for doctors, appear repeatedly as the key registration/licensing gatekeepers discussed.
- Career-journey style: questions follow a natural job-seeker journey — from finding an agency/vacancy, through visa/work-permit and professional registration, to interview prep, CV writing, salary, relocation, and career growth once employed.
- No duplicate records or questions: all 97
idvalues and all 97 question texts are unique. - No malformed JSON lines: every line parses successfully with a consistent flat schema.
- Sequential, gap-free IDs:
global_healthcare_jobs_001throughglobal_healthcare_jobs_097, matching the reported row count exactly.
8. Suggested Use Cases
- Fine-tuning or evaluating Nepali-language LLMs for career guidance / overseas job-migration counseling for healthcare workers.
- Domain-adaptation for Nepali-language recruitment/immigration chatbots targeting healthcare professionals seeking UK employment.
- Benchmarking short, direct factual-answer generation in Nepali (contrast with the longer explanatory paragraphs in the Vitiligo dataset).
- Supporting Nepali diaspora / migrant-worker information services, given Nepal's large healthcare-worker migration pipeline to the UK and other countries.
- As a contrast set against the Fight Vitiligo dataset to study answer-length and behavior-definition variation across "real" (non-synthetic) Nepali FAQ datasets from different domains.
9. Known Limitations
- Small dataset size: only 97 records — useful for narrow fine-tuning/evaluation on this specific FAQ topic, but too small alone for broad instruction tuning.
- Single domain, single behavior: entirely focused on UK NHS job-seeking FAQ with one behavior type (
जानकारीमूलक उत्तर दिने) — should be combined with other datasets for general-purpose or multi-domain instruction tuning. - UK-specific content: institutional names, processes, and regulatory bodies (NMC, GMC, NHS-specific terminology) are UK-specific and may not generalize to healthcare job-seeking guidance for other countries.
- No length/complexity metadata: unlike the Vitiligo dataset, this dataset provides no
question_length,response_length,maximum_response_sentences,minimum_reasoning_dimensions, orminimum_complexity_scorefields, limiting fine-grained filtering by answer length/complexity. - No explicit source URLs: the
urlfield is empty for every record, so individual FAQ entries cannot be traced back to a specific original web page. - Time-sensitivity: visa rules, registration processes, and NHS recruitment practices change over time; the factual accuracy of specific procedural claims (e.g. current sponsor licence rules) should be independently verified before use in a production advisory context.
