CoolFace
Datasetpublic

faisal4590aziz/bangla-health-related-paraphrased-dataset

Dataset Card for "BanglaHealthParaphrase" BanglaHealthParaphrase is a Bengali paraphrasing dataset specifically curated for the health domain. It contains over 200,000 sentence pairs, where each pair consists of an original Bengali sentence and its paraphrased version. The dataset was created through a multi-step pipeline involving extraction of health-related content from Bengali news sources, English pivot-based paraphrasing, and back-translation to ensure linguistic… See the full description on the dataset page: https://huggingface.co/datasets/faisal4590aziz/bangla-health-related-paraphrased-dataset.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
2likes24downloads
Dataset Card

Dataset Card for "BanglaHealthParaphrase"

<!-- Provide a quick summary of the dataset. -->

BanglaHealthParaphrase is a Bengali paraphrasing dataset specifically curated for the health domain. It contains over 200,000 sentence pairs, where each pair consists of an original Bengali sentence and its paraphrased version. The dataset was created through a multi-step pipeline involving extraction of health-related content from Bengali news sources, English pivot-based paraphrasing, and back-translation to ensure linguistic diversity while preserving semantic integrity. A custom python script was developed to achieve this task. This dataset aims to support research in Bengali paraphrase generation, text simplification, chatbot development, and domain-specific NLP applications.

Dataset Description

FieldDescription
LanguageBengali (Bangla)
DomainHealth
Size200,000 sentence pairs
FormatCSV with fields: sl, id, source_sentence, paraphrase_sentence
LicenseCC BY 4.0

Each entry contains:

  • —SL: SL represents serial number of the sentence pair.
  • —ID: ID represents unique identifier for each sentence pair.
  • —source_sentence: A Bengali sentence from a health-related article.
  • —paraphrased_sentence: A meaning-preserving paraphrase of the source sentence.

🛠️ Construction Methodology

The dataset was developed through the following steps:

  1. 1.Data Collection: Health-related sentences were scraped from leading Bengali online newspapers and portals.
  2. 2.Preprocessing: Cleaned and normalized sentences using custom text-cleaning scripts.
  3. 3.Pivot Translation Method:
  4. 4.Translated original Bengali sentences into English.
  5. 5.Generated English paraphrases using a T5-based paraphraser.
  6. 6.Back-translated English paraphrases to Bengali using NMT models.
  7. 7.Filtering: Applied semantic similarity filtering and sentence length constraints.
  8. 8.Manual Validation: A sample subset was manually reviewed by 100 peers to ensure quality and diversity.

✨ Key Features

  • —Focused entirely on the Healthcare Domain.
  • —High-quality paraphrases generated using a robust multilingual pivoting approach.
  • —Suitable for use in Bengali NLP tasks such as paraphrase generation, low-resource model training, and chatbot development.

📰 Related Publication

This dataset is published in Data in Brief (Elsevier):

BanglaHealth: A Bengali paraphrase Dataset on Health Domain Faisal Ibn Aziz, Dr. Mohammad Nazrul Islam, 2025

Dataset Sources

Detailed in the paper

Licensing Information

Contents of this repository are restricted to only non-commercial research purposes under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (CC BY-NC-SA 4.0). Copyright of the dataset contents belongs to the original copyright holders.

Data Instances

Sample data format.

{
  "sl": 28,
  "id": 28,
  "source_sentence": "ডেঙ্গু হেমোরেজিক ফিভারে মূলত রোগীর রক্তনালীগুলোর দেয়ালে যে ছোট ছোট ছিদ্র থাকে, সেগুলো বড় হয়ে যায়।",
  "paraphrased_sentence": "ডেঙ্গু হেমোরেজিক ফিভারে রোগীর রক্তনালীর দেয়ালে ছোট ছোট ছিদ্র বড় হয়ে যায়।"
}

Loading the dataset

python
from datasets import load_dataset
dataset = load_dataset("faisal4590aziz/bangla-health-related-paraphrased-dataset")

Citation

If you are using this dataset, please cite this in your research.

@article{AZIZ2025111699,
  title = {BanglaHealth: A Bengali paraphrase Dataset on Health Domain},
  journal = {Data in Brief},
  pages = {111699},
  year = {2025},
  issn = {2352-3409},
  doi = {https://doi.org/10.1016/j.dib.2025.111699},
  url = {https://www.sciencedirect.com/science/article/pii/S2352340925004299},
  author = {Faisal Ibn Aziz and Muhammad Nazrul Islam},
  keywords = {Natural Language Processing (NLP), Paraphrasing, Bengali Paraphrasing, Bengali Language, Health Domain},
}