datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bioleaflets-biomedical-ner
Dataset Card for BioLeaflets Dataset
Dataset Summary
BioLeaflets is a biomedical dataset for Data2Text generation. It is a corpus of 1,336 package leaflets of medicines authorised in Europe, which were obtained by scraping the European Medicines Agency (EMA) website.
Package leaflets are included in the packaging of medicinal products and contain information to help patients use the product safely and appropriately.
This dataset comprises the large majority (∼ 90%) of… See the full description on the dataset page: https://huggingface.co/datasets/ruslan/bioleaflets-biomedical-ner.Tutorbot-Spock-Bio-DatasetMock conversations between a student and a tutor to train a chatbot for educational purposes as suggested in the paper
CLASS Meet SPOCK: An Education Tutoring Chatbot based on Learning Science Principles.
Dataset generated from OpenStax Biology 2e textbook.
Problem, Subproblem, Hints, and Feedback is generated using the prompt.
Mock Conversations is generated using the prompt.
For any queries, contact Shashank Sonkar (ss164 AT rice dot edu)
If you use this model, please cite:
CLASS Meet… See the full description on the dataset page: https://huggingface.co/datasets/luffycodes/Tutorbot-Spock-Bio-Dataset.NCERT_Biology_11thNCERT_Biology_12thGMASS-probe-set-v1.0
MediSafe-GH: A Clinical Safety Screen for Medical AI Assistants in Ghanaian Languages
Project Summary
We are developing G-MASS (Ghana Medical AI Safety Screen), an open-source, reusable evaluation protocol that tests whether AI health assistants give safe responses (not just accurate ones) to medical queries posed in standard English, Twi, and Ghanaian English, for use by health AI developers and clinical technology researchers.
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/BioinstLab/GMASS-probe-set-v1.0.clinical-quad-pk-sampling-window-deviation-bioanalytical-variance-dose-adjustment-interim-v0.1Clarus Clinical Quad Coupling PK Integrity v0.1
PurposeDetect PK integrity distortion driven by four interacting nodes.
Quad nodes
Sampling window deviation
Bioanalytical or stability variance
Dose adjustment decisions
Governance interim or submission timing
InputOne vignette.
OutputStrict JSON only.
Required keys
pk_integrity_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence
Filesdata/train.csvdata/test.csvscorer.py
Run scoringCreate… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-pk-sampling-window-deviation-bioanalytical-variance-dose-adjustment-interim-v0.1.illicit-bio-multi-turn
Illicit Bio Multi-Turn Conversations
Multi-turn adversarial conversations that successfully elicited harmful bio-safety content from AI models. This sample dataset contains 5 conversations (57 turns) covering bioweapons and related threats.
Dataset Statistics
Metric
Value
Conversations
5
Total Turns
57
Avg Turns/Conv
11.4
Harm Categories
3
Harm Categories
Category
Turns
Description
Bioweapons
34
Information about biological… See the full description on the dataset page: https://huggingface.co/datasets/GoJulyAI/illicit-bio-multi-turn.rag-bioask-plDerived from rag-datasets/mini-bioasq.
This is a subset of the above dataset translated to Polish using DeepL. It should help you in finetuning LLMs for RAG purposes.
bioactives-naturals-smiles-molgen
Valid Bioactives and Natural Product SMILES
~2.7M valid SMILES built and curated from ChemBL34 (Zdrazil et al. 2023), COCONUTDB (Sorokina et al. 2021), and Supernatural3 (Gallo et al. 2023) dataset.
Curated by: gbyuvd
References
BibTeX
COCONUTDB
@article{sorokina2021coconut,
title={COCONUT online: Collection of Open Natural Products database},
author={Sorokina, Maria and Merseburger, Peter and Rajan, Kohulan and Yirik, Mehmet Aziz and Steinbeck… See the full description on the dataset page: https://huggingface.co/datasets/gbyuvd/bioactives-naturals-smiles-molgen.ligase_bioremedation_sequencesNCERT_Biology_11thRoomly-Student-Bios-Multimodal
Roomly: Multimodal Roommate Matching Dataset
🎯 Problem Statement
Finding a roommate is often reduced to dry filters like "budget" and "location". Roomly aims to revolutionize this by focusing on personality, lifestyle, and visual preferences. This dataset provides synthetic student profiles and their ideal room environments.
📊 Exploratory Data Analysis (EDA)
1. User Persona Distribution
Our dataset contains a balanced mix of different student… See the full description on the dataset page: https://huggingface.co/datasets/Orib24/Roomly-Student-Bios-Multimodal.
