datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
meerkat-instructionsThis repository provides the instruction tuning data used to train our medical language model, Meerkat, along with descriptions.
For more information, please refer to the paper below. Our models can be downloaded from the official model repository.
📄 Paper: Small Language Models Learn Enhanced Reasoning Skills from Medical Textbooks
Dataset Statistics
Table: Statistics of our instruction-tuning datasets“# Examples” denotes the number of training examples for each dataset.
†… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/meerkat-instructions.oc20-s2ef-uma-embeddingsMedLFQAOriginal dataset introduced by Jeong et al. in OLAPH: Improving Factuality in Biomedical Long-form Question Answering
Citation information:
@misc{jeong2024olaph,
title={OLAPH: Improving Factuality in Biomedical Long-form Question Answering},
author={Minbyul Jeong and Hyeon Hwang and Chanwoong Yoon and Taewhoo Lee and Jaewoo Kang},
year={2024},
eprint={2405.12701},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
llama-3.1-medprm-reward-training-set
Med-PRM-Reward (Version 1.0)
🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-training-set.ToxReason
ToxReason
🚀 Accepted at ACL 2026 Findings
ToxReason is a benchmark dataset for mechanistic chemical toxicity reasoning based on Adverse Outcome Pathways (AOPs).
The dataset is designed to evaluate whether large language models can generate biologically interpretable toxicity reasoning trajectories that connect molecular structures, Molecular Initiating Events (MIEs), pathway perturbations, and organ-level adverse outcomes.
Dataset Overview
ToxReason consists of… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/ToxReason.planned_diffusion_reasoningETHICGithub : https://github.com/dmis-lab/ETHICArxiv : https://arxiv.org/abs/2410.16848
llama-3.1-medprm-reward-raw-training-setpdv2-dataoc20-s2ef-vae-embeddings-devTemporalHead
[ACL 2025] Does Time Have Its Place? Temporal Heads: Where Language Models Recall Time-specific Information
This repository contains two separate subsets of data (configs):
Temporal: JSON files in Temporal that include temporal knowledge.
Invariant: JSON files in Invariant that describe time-invariant knowledge based on LRE.
Each subset has its own schema. By defining them as two configs in the YAML header above, Hugging Face’s Dataset Viewer will show “Temporal” and… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/TemporalHead.D_mis72_run0_gen8_WXS_doc1000_synt64_lr1e-04_acm_FRESHD_mis72_run0_gen7_WXS_doc1000_synt64_lr1e-04_acm_FRESHD_mis72_run0_gen1_WXS_doc1000_synt64_lr1e-04_acm_MPPessential-web-subsample
essential-web-subsample
A small, category-balanced subsample of Essential-Web v1.0 for studying how generation behavior (e.g. diffusion-LM parallelism) varies with data type.
900 documents = 100 per category across 9 categories, mapped from Essential-Web's document_type_v1.primary.label:
category
Essential-Web document_type_v1
code
Code/Software
papers
Academic/Research
encyclopedic
Reference/Encyclopedic/Educational
legal
Legal/Regulatory
literary… See the full description on the dataset page: https://huggingface.co/datasets/dmisrael/essential-web-subsample.oc20-s2ef-vae-embeddings-testoc20-s2ef-uma-embeddings-devoc20-s2ef-vae-embeddings-test4pdv2_testD_mis_run3_gen8_WXS_doc1000_synt64_lr1e-04_acm_SYNLASTD_mis72_run0_gen2_WXS_doc1000_synt64_lr1e-04_acm_MPPpdv2_test_1.2D_mis72_run0_gen4_WXS_doc1000_synt64_lr1e-04_acm_FRESHD_mis73_run0_gen3_WXS_doc1000_synt64_lr1e-04_acm_FRESHD_mis73_run0_gen4_WXS_doc1000_synt64_lr1e-04_acm_FRESHpdv2_test_1.1D_mis72_run0_gen10_WXS_doc1000_synt64_lr1e-04_acm_FRESHoc20-s2ef-vae-embeddings-test3D_mis72_run0_gen3_WXS_doc1000_synt64_lr1e-04_acm_FRESHD_mis72_run0_gen0_WXS_doc1000_synt64_lr1e-04_acm_MPP
