Bibek-Poudel/ENG_NEP_MED_PARALLEL
Dataset Card for Dataset Name This dataset aims to be a state of art Nepali English Parallel translation in Medical Domain. Further on this dataset will be updated with more correct translations. Dataset Details Dataset Description Curated by: Bibek Poudel Language(s) (NLP): Nepali, English Dataset Sources [optional] Repository: https://www.kaggle.com/datasets/rxnach/nepali-health-forum-corpus-questions-and-answers… See the full description on the dataset page: https://huggingface.co/datasets/Bibek-Poudel/ENG_NEP_MED_PARALLEL.
Dataset Card for Dataset Name
<!-- Provide a quick summary of the dataset. -->
This dataset aims to be a state of art Nepali English Parallel translation in Medical Domain. Further on this dataset will be updated with more correct translations.
Dataset Details
Dataset Description
<!-- Provide a longer summary of what this dataset is. -->
- Curated by: Bibek Poudel
- Language(s) (NLP): Nepali, English
Dataset Sources [optional]
<!-- Provide the basic links for the dataset. -->
- Repository: https://www.kaggle.com/datasets/rxnach/nepali-health-forum-corpus-questions-and-answers
- Repository: https://www.kaggle.com/datasets/poudelsujan03/pregnancy-related-question-answer-nepali-dataset
- Website: https://ekantiput.com
Uses
Machine Translation (MT) Train and evaluate NE→EN and EN→NE translation models for low-resource languages. Domain-Adaptive MT Improve existing models (e.g. mT5, NLLB, MarianMT) with domain-specific fine-tuning in medicine and education. Healthcare NLP Build or benchmark AI systems that understand Nepali-language medical content — useful for public health chatbots, translation of health records, or instructional text generation. Low-Resource Language Research Offers a rare parallel corpus in Nepali, especially in specialized domains. Cross-Lingual Transfer Learning Improve downstream tasks such as summarization, classification, and question answering across languages. Curriculum Learning / Multi-Stage Training Supports phased fine-tuning where models start with general domain and specialize in medical content.
Direct Use
NLP LLM Fine Tuning
[More Information Needed]
Dataset Structure
Dataset has been divided into Train Val and Test in the ratio of 80,10,10
[More Information Needed]
Dataset Creation
9 Apr 2025
Recommendations
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
Users should be made aware of the risks, biases and limitations of the dataset. More information needed for further recommendations.
More Information [optional]
poudelbibek38@gmail.com
Dataset Card Authors [optional]
Bibek Poudel
Dataset Card Contact
poudelbibek38@gmail.com
