CoolFace
Datasetpublic

sabin1234/Global_Healthcare_Jobs_FAQ

Global Healthcare Jobs FAQ — UK NHS Careers Nepali Q&A Dataset 1. Overview This dataset is a Nepali-language (Devanagari script) collection of question–answer pairs in ShareGPT format, covering frequently asked questions about working in the UK's National Health Service (NHS) — aimed specifically at international/foreign healthcare workers (nurses, doctors, dentists, and other clinicians) seeking employment, registration, visas, and career guidance in the United… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Global_Healthcare_Jobs_FAQ.

sourceHugging Faceapache-2.0updated 15d agoView on Hugging Face
0likes40downloads
Dataset Card

Global Healthcare Jobs FAQ — UK NHS Careers Nepali Q&A Dataset

1. Overview

This dataset is a Nepali-language (Devanagari script) collection of question–answer pairs in ShareGPT format, covering frequently asked questions about working in the UK's National Health Service (NHS) — aimed specifically at international/foreign healthcare workers (nurses, doctors, dentists, and other clinicians) seeking employment, registration, visas, and career guidance in the United Kingdom's healthcare sector.

Each record is a single-turn conversation where a human asks a practical, career-related question in Nepali about UK NHS job-seeking, and the assistant ("gpt") gives a short, factual, informative answer in Nepali.

PropertyValue
File nameglobal_healthcare_jobs_faq_nepali__1_.jsonl
File formatJSON Lines (.jsonl) — one JSON object per line
Total records (rows)97
Conversation formatShareGPT (conversations list with from/value pairs)
Turns per conversation2 (1 human turn + 1 gpt turn) — 100% of records
DomainHealth / Careers — UK NHS jobs for international healthcare workers
Content languageNepali (Devanagari script) — 100% verified
Data origin"Global Healthcare Jobs" FAQ content (real-world, not synthetically generated)
License (per-record)Apache-2.0 (permissive)
Duplicate IDs found0 (all 97 id values unique)
Duplicate questions found0
Malformed / unparseable lines0

2. File & Record Structure

This dataset uses the same flat schema style seen in the Fight Vitiligo FAQ dataset — most provenance metadata sits at the top level, with fine-grained generation metadata nested in a stringified metadata_json field.

KeyTypeDescription
idstringUnique record identifier, format global_healthcare_jobs_0XX (e.g. global_healthcare_jobs_001).
conversationsarrayShareGPT-style conversation turns. Always exactly 2 turns.
conversations[0].fromstringAlways "human" — the question turn.
conversations[0].valuestringThe Nepali-language question about UK NHS jobs/careers.
conversations[1].fromstringAlways "gpt" — the answer turn.
conversations[1].valuestringA short, factual, Nepali-language answer.
sourcestringName of the originating content source.
source_namestringMachine-style source identifier (snake_case).
source_repostringRepository/collection name.
source_configstringConfig label.
source_splitstringData split label.
source_revisionstringVersion tag.
source_row_idstringRow-level traceability ID (identical to top-level id).
languagestringISO language code.
language_codestringISO 639-3 code.
scriptstringScript tag.
licensestringLicense of the content.
license_tierstringLicense permissiveness bucket.
task_typestringNLP task type this record represents.
generation_typestringHow the record was produced.
conditionstringGeneration condition tag.
urlstringSource URL reference (always empty in this file).
metadata_jsonstringStringified JSON object with fine-grained generation metadata (parsed in Section 3).

2.1 Top-level provenance fields (all constant across the dataset)

FieldObserved value%
sourceGlobal Healthcare Jobs100%
source_nameglobal_healthcare_jobs_faq100%
source_repoGlobal Healthcare Jobs100%
source_configdefault100%
source_splittrain100%
source_revisionv1100%
languagene100%
language_codenpi100%
scriptDeva (Devanagari)100%
licenseApache-2.0100%
license_tierpermissive100%
task_typeinstruction-following100%
generation_typereal100%
conditionreal100%
url"" (empty)100%

As with the Fight Vitiligo dataset, this dataset is tagged `generation_type: real` / `condition: real`, indicating it is derived from genuine FAQ content rather than being purely template-generated.

2.2 metadata_json sub-fields (parsed)

FieldDescriptionObserved value(s)
generation_domainTop-level subject domain.स्वास्थ्य ("Health") — 100%
generation_categorySub-topic category.संयुक्त अधिराज्य राष्ट्रिय स्वास्थ्य सेवा जागिर सम्बन्धी प्रायः सोधिने प्रश्न ("Frequently asked questions about UK National Health Service jobs") — 100%
generation_sub_domainSpecific FAQ theme.राष्ट्रिय स्वास्थ्य सेवा जागिर, प्रवेशाज्ञा र दर्ता ("NHS jobs, work permits, and registration") — 100%
behaviorExpected assistant behavior.जानकारीमूलक उत्तर दिने ("Provide an informative answer") — 100%
behavior_definitionFull definition of the expected behavior."मूल प्रश्नको अर्थ र मूल उत्तरको तथ्य कायम राख्दै शुद्ध नेपालीमा रूपान्तरण गरिएको उत्तर दिने।" ("Give an answer translated into pure Nepali while preserving the meaning of the original question and the facts of the original answer.") — 100%
question_typeFormat classification.तथ्यमा आधारित प्रश्न ("Fact-based question") — 100%
content_languageDeclared content language.नेपाली ("Nepali") — 100%, accurate
content_scriptDeclared content script.देवनागरी ("Devanagari") — 100%, accurate
english_content_allowedWhether English content is permitted.False — 100%

Notable schema difference: unlike the Fight Vitiligo dataset, this dataset's metadata_json does not include question_length, response_length, maximum_response_sentences, minimum_reasoning_dimensions, or minimum_complexity_score fields — those length/complexity constraints are simply absent here, suggesting this dataset was generated/curated under a lighter-weight metadata schema (likely an earlier or simpler pipeline version, or a straight translation task — note the behavior definition explicitly says the answer is "translated into pure Nepali... preserving the original question's meaning and the original answer's facts," implying the source FAQ content was originally in English and has been translated to Nepali, unlike the Vitiligo dataset's more generic "give a fact-based answer" framing).


3. Question Distribution

3.1 By domain/category/sub-domain

The dataset is single-domain and single-topic by design — every record belongs to the same taxonomy path:

LevelValueCount%
generation_domainस्वास्थ्य (Health)97100%
generation_categoryUK NHS jobs FAQ97100%
generation_sub_domainNHS jobs, work permits & registration97100%
question_typeतथ्यमा आधारित प्रश्न (Fact-based question)97100%

All questions are open-ended and explanatory (no MCQ format) — the assistant is expected to provide a direct, factual, career-advice-style answer.

3.2 Thematic sub-topic breakdown (derived from question content)

The formal taxonomy is uniform, but the actual question content clusters around several recognizable career-related sub-themes. The table below is a keyword-based approximate clustering (a question may match more than one theme, so counts can overlap):

Sub-theme (approx.)Example keyword(s) matchedQuestions matched% of dataset
Registration / licensing (NMC, GMC, etc.)दर्ता / लाइसेन्स / रजिस्ट्रेसन99.3%
Qualification / experience / equivalencyयोग्यता / अनुभव / प्रमाणपत्र99.3%
International / overseas recruitmentविदेशी / अन्तर्राष्ट्रिय77.2%
Visa / work permit / immigrationभिसा / प्रवेशाज्ञा / आप्रवासन77.2%
Temporary / locum positionsअस्थायी / स्थानापन्न / लोकम66.2%
CV / resume / job applicationसिभी / रिज्युमे / आवेदन66.2%
Job search / vacanciesजागिर खोज / रिक्त / अवसर44.1%
Recruitment agenciesएजेन्सी / भर्ती44.1%
Interviewsअन्तर्वार्ता44.1%
Salary / payतलब / पारिश्रमिक / भुक्तानी44.1%
Training / inductionतालिम / अभिमुखीकरण44.1%
Relocation / accommodationसरुवा / बसोबास / आवास22.1%
Language requirement (English/IELTS/OET)अंग्रेजी / IELTS / OET11.0%
Other / unmatched by keyword list4546.4%

52 of 97 questions (53.6%) matched at least one of the above keyword groups; the remainder cover related but more loosely-worded aspects of the UK healthcare job-search journey (e.g. career progression, professional networking, specific professional-council processes like changing a nursing PIN, sponsor licences, medical specialty registration, competitiveness of roles). Registration/licensing and qualification/experience recognition are the most prominent explicit themes, reflecting the central pain point for foreign-trained clinicians: getting credentials recognized by UK regulatory bodies (e.g. the Nursing and Midwifery Council for nurses, the General Medical Council for doctors) before they can legally practice.

3.3 Length characteristics

MetricMinMedianMeanMax
Question — character length427878.6111
Question — word count71111.618
Answer — character length46173173.5292
Answer — word count72323.439
Answer — sentence count (approx.)111.64

Answers here are noticeably shorter than the Fight Vitiligo dataset's (mean 173.5 characters / 23.4 words here, vs. mean 305.2 characters / 45.6 words for Vitiligo) — typically 1–2 concise sentences giving a direct, practical answer rather than a multi-sentence explanatory paragraph. This fits a "quick FAQ / career-tips" style rather than in-depth medical education content.

3.4 Answer format

Answers are free-form, direct, and practical — short factual statements or brief pieces of guidance (e.g. naming example recruitment agencies, confirming a process exists, or giving a concise yes/no-plus-explanation). No lettered options or MCQ structure is present anywhere in the dataset.


4. Behavior Distribution

The dataset targets a single, uniform assistant behavior:

BehaviorBehavior definitionCount%
जानकारीमूलक उत्तर दिने ("Provide an informative answer")"Give an answer translated into pure Nepali while preserving the meaning of the original question and the facts of the original answer."97100%

4.1 Behavior-related generation notes

Unlike the two previously reviewed datasets, this dataset's metadata_json does not define numeric constraints such as maximum_response_sentences, minimum_reasoning_dimensions, or minimum_complexity_score — these fields are simply not present in any record. The only explicit behavioral constraint is:

ConstraintValueMeaning
english_content_allowedFalseNo English content permitted anywhere in the record — output must be pure Nepali.

The behavior definition's phrasing ("preserving the meaning of the original question and the facts of the original answer... translated into pure Nepali") strongly suggests this dataset was produced by translating an existing English-language NHS-careers FAQ resource into Nepali, rather than being generated from scratch or templated from structured data (contrast with the BFI banking dataset, which was templated from a structured registry, or the Vitiligo dataset, which was framed as direct fact-based Q&A generation).

4.2 Task type / generation type

FieldValueCount%
task_typeinstruction-following97100%
generation_typereal97100%
conditionreal97100%

5. Provenance & Licensing Distribution

FieldValueCount%
sourceGlobal Healthcare Jobs97100%
source_nameglobal_healthcare_jobs_faq97100%
source_repoGlobal Healthcare Jobs97100%
source_splittrain97100%
source_revisionv197100%
licenseApache-2.097100%
license_tierpermissive97100%
  • Single source: all 97 records trace back to one content source ("Global Healthcare Jobs" — a UK-focused healthcare recruitment/career-advice publisher).
  • License: Apache-2.0, same permissive license tier as the Fight Vitiligo dataset (differs from the CC0-1.0 used in the BFI banking dataset). Apache-2.0 permits reuse and modification but expects attribution/notice preservation.
  • Full traceability: each record's source_row_id matches its id directly (global_healthcare_jobs_001global_healthcare_jobs_097), giving a simple, sequential trace back to source ordering.
  • No URL field populated: every record's url field is an empty string, so there is no direct link back to a specific original FAQ page per record.

6. Language & Script Notes

This dataset's language metadata is accurate and consistent with its actual content, matching the pattern seen in the Fight Vitiligo dataset:

FieldDeclared valueActual content observed
languageneNepali ✅
language_codenpiNepali (ISO 639-3: npi) ✅
scriptDevaDevanagari ✅
metadata_json.content_languageनेपाली (Nepali)Nepali ✅
metadata_json.content_scriptदेवनागरी (Devanagari)Devanagari ✅
metadata_json.english_content_allowedFalseNo English text found ✅

Verified finding: A full script scan of all 97 records confirms 100% of both questions and answers are written entirely in Devanagari script — zero Latin-only text in either turn. Note, however, that several UK-specific institutional terms are transliterated into Nepali phonetically rather than translated (e.g. "राष्ट्रिय स्वास्थ्य सेवा" for "National Health Service," "नर्सिङ तथा मिडवाइफरी परिषद" for "Nursing and Midwifery Council," "सामान्य चिकित्सा परिषद" for "General Medical Council") — this is a translation/localization choice, not a script violation, since the surface text remains fully Devanagari.


7. Content Coverage Notes

  • Single career-domain focus: every question concerns working in the UK NHS as a foreign/international healthcare professional — spanning nurses, doctors, dentists, and other clinical roles.
  • Regulatory bodies referenced: the Nursing and Midwifery Council (NMC) for nurses, and the General Medical Council (GMC) for doctors, appear repeatedly as the key registration/licensing gatekeepers discussed.
  • Career-journey style: questions follow a natural job-seeker journey — from finding an agency/vacancy, through visa/work-permit and professional registration, to interview prep, CV writing, salary, relocation, and career growth once employed.
  • No duplicate records or questions: all 97 id values and all 97 question texts are unique.
  • No malformed JSON lines: every line parses successfully with a consistent flat schema.
  • Sequential, gap-free IDs: global_healthcare_jobs_001 through global_healthcare_jobs_097, matching the reported row count exactly.

8. Suggested Use Cases

  • Fine-tuning or evaluating Nepali-language LLMs for career guidance / overseas job-migration counseling for healthcare workers.
  • Domain-adaptation for Nepali-language recruitment/immigration chatbots targeting healthcare professionals seeking UK employment.
  • Benchmarking short, direct factual-answer generation in Nepali (contrast with the longer explanatory paragraphs in the Vitiligo dataset).
  • Supporting Nepali diaspora / migrant-worker information services, given Nepal's large healthcare-worker migration pipeline to the UK and other countries.
  • As a contrast set against the Fight Vitiligo dataset to study answer-length and behavior-definition variation across "real" (non-synthetic) Nepali FAQ datasets from different domains.

9. Known Limitations

  • Small dataset size: only 97 records — useful for narrow fine-tuning/evaluation on this specific FAQ topic, but too small alone for broad instruction tuning.
  • Single domain, single behavior: entirely focused on UK NHS job-seeking FAQ with one behavior type (जानकारीमूलक उत्तर दिने) — should be combined with other datasets for general-purpose or multi-domain instruction tuning.
  • UK-specific content: institutional names, processes, and regulatory bodies (NMC, GMC, NHS-specific terminology) are UK-specific and may not generalize to healthcare job-seeking guidance for other countries.
  • No length/complexity metadata: unlike the Vitiligo dataset, this dataset provides no question_length, response_length, maximum_response_sentences, minimum_reasoning_dimensions, or minimum_complexity_score fields, limiting fine-grained filtering by answer length/complexity.
  • No explicit source URLs: the url field is empty for every record, so individual FAQ entries cannot be traced back to a specific original web page.
  • Time-sensitivity: visa rules, registration processes, and NHS recruitment practices change over time; the factual accuracy of specific procedural claims (e.g. current sponsor licence rules) should be independently verified before use in a production advisory context.