datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
thai_buddhist_studies_exam
Thai Buddhist Studies Examination (Nak Tham)
This repository contains multiple-choice questions from the Thai Buddhist Studies
(Nak Tham) examination (2020, 2022, 2023). This dataset can be used for a benchmark for evaluating Large Language Models'
understanding of Thai Buddhist concepts and teachings.
Dataset Statistics
Year
Number of Multiple Choice Questions
2020
1,350
2022
1,400
2023
1,350
Phra Udom thought on the exam: We have reviewed the Nak… See the full description on the dataset page: https://huggingface.co/datasets/biodatlab/thai_buddhist_studies_exam.BiomixQA
BiomixQA Dataset
Overview
BiomixQA is a curated biomedical question-answering dataset comprising two distinct components:
Multiple Choice Questions (MCQ)
True/False Questions
This dataset has been utilized to validate the Knowledge Graph based Retrieval-Augmented Generation (KG-RAG) framework across different Large Language Models (LLMs). The diverse nature of questions in this dataset, spanning multiple choice and true/false formats, along with its coverage of various… See the full description on the dataset page: https://huggingface.co/datasets/kg-rag/BiomixQA.predator-biomedical
PREDATOR Biomedical Dataset
50K+ curated biomedical abstracts from PubMed/EuropePMC
Curated biomedical research data extracted from PubMed, EuropePMC, and clinical trial databases. Each entry includes DOI/PMID, title, source, domain classification, commercial value score, and quality assessment.
Fields
Column
Description
id
DOI or PMID identifier
title
Article title or abstract summary
source
Data source (EuropePMC, PubMed, ClinicalTrials.gov)… See the full description on the dataset page: https://huggingface.co/datasets/iservice/predator-biomedical.BiomedKeyRefine
BiomedKeyRefine: Refined Biomedical Keyword Extraction Benchmark
Version 1.0Last updated: February 2026Author: Christian Hoang (@christhoang04)License: MIT
🌟 Dataset Summary
BiomedKeyRefine is a high-quality, refined subset of the PubMedAKE benchmark (CIKM 2022), specifically curated for biomedical keyword extraction research.
The original PubMedAKE contains 843k+ articles with author-assigned keywords (extractive + abstractive). This version:
Filters for… See the full description on the dataset page: https://huggingface.co/datasets/christian-hoang-04/BiomedKeyRefine.GMASS-probe-set-v1.0
MediSafe-GH: A Clinical Safety Screen for Medical AI Assistants in Ghanaian Languages
Project Summary
We are developing G-MASS (Ghana Medical AI Safety Screen), an open-source, reusable evaluation protocol that tests whether AI health assistants give safe responses (not just accurate ones) to medical queries posed in standard English, Twi, and Ghanaian English, for use by health AI developers and clinical technology researchers.
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/BioinstLab/GMASS-probe-set-v1.0.biomedical_cpgQA
Dataset Card for the Biomedical Domain
Dataset Summary
This dataset was obtain through github (https://github.com/mmahbub/cpgQA/blob/main/dataset/cpgQA-v1.0.csv?plain=1) to Huggin Face for easier access while fine tuning.
Languages
English (en)
Dataset Structure
The dataset is in a CSV format, with each row representing a single review. The following columns are included:
Title: Categorises the QA.
Context: Gives a context of the QA.
Question: The… See the full description on the dataset page: https://huggingface.co/datasets/chloecchng/biomedical_cpgQA.NCERT_Biology_11thNCERT_Biology_12thBiomixQA
BiomixQA Dataset
Overview
BiomixQA is a curated biomedical question-answering dataset comprising two distinct components:
Multiple Choice Questions (MCQ)
True/False Questions
This dataset has been utilized to validate the Knowledge Graph based Retrieval-Augmented Generation (KG-RAG) framework across different Large Language Models (LLMs). The diverse nature of questions in this dataset, spanning multiple choice and true/false formats, along with its coverage of various… See the full description on the dataset page: https://huggingface.co/datasets/shakthish/BiomixQA.biologyhelper
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/dddjjjppp/biologyhelper.NCERT_Biology_11thMedRisk-BioMistral
Reference:
"A Question-Entailment Approach to Question Answering". Asma Ben Abacha and Dina Demner-Fushman. BMC Bioinformatics, 2019.
