CoolFace
Datasetpublic

Bibek-Poudel/ENG_NEP_MED_PARALLEL

Dataset Card for Dataset Name This dataset aims to be a state of art Nepali English Parallel translation in Medical Domain. Further on this dataset will be updated with more correct translations. Dataset Details Dataset Description Curated by: Bibek Poudel Language(s) (NLP): Nepali, English Dataset Sources [optional] Repository: https://www.kaggle.com/datasets/rxnach/nepali-health-forum-corpus-questions-and-answers… See the full description on the dataset page: https://huggingface.co/datasets/Bibek-Poudel/ENG_NEP_MED_PARALLEL.

sourceHugging Faceupdated 1y agoView on Hugging Face
1likes70downloads
Dataset Card

Dataset Card for Dataset Name

<!-- Provide a quick summary of the dataset. -->

This dataset aims to be a state of art Nepali English Parallel translation in Medical Domain. Further on this dataset will be updated with more correct translations.

Dataset Details

Dataset Description

<!-- Provide a longer summary of what this dataset is. -->

  • Curated by: Bibek Poudel
  • Language(s) (NLP): Nepali, English

Dataset Sources [optional]

<!-- Provide the basic links for the dataset. -->

  • Repository: https://www.kaggle.com/datasets/rxnach/nepali-health-forum-corpus-questions-and-answers
  • Repository: https://www.kaggle.com/datasets/poudelsujan03/pregnancy-related-question-answer-nepali-dataset
  • Website: https://ekantiput.com

Uses

Machine Translation (MT) Train and evaluate NE→EN and EN→NE translation models for low-resource languages. Domain-Adaptive MT Improve existing models (e.g. mT5, NLLB, MarianMT) with domain-specific fine-tuning in medicine and education. Healthcare NLP Build or benchmark AI systems that understand Nepali-language medical content — useful for public health chatbots, translation of health records, or instructional text generation. Low-Resource Language Research Offers a rare parallel corpus in Nepali, especially in specialized domains. Cross-Lingual Transfer Learning Improve downstream tasks such as summarization, classification, and question answering across languages. Curriculum Learning / Multi-Stage Training Supports phased fine-tuning where models start with general domain and specialize in medical content.

Direct Use

NLP LLM Fine Tuning

[More Information Needed]

Dataset Structure

Dataset has been divided into Train Val and Test in the ratio of 80,10,10

[More Information Needed]

Dataset Creation

9 Apr 2025

Recommendations

<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->

Users should be made aware of the risks, biases and limitations of the dataset. More information needed for further recommendations.

More Information [optional]

poudelbibek38@gmail.com

Dataset Card Authors [optional]

Bibek Poudel

Dataset Card Contact

poudelbibek38@gmail.com