CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dmis-lab /meerkat-instructionsThis repository provides the instruction tuning data used to train our medical language model, Meerkat, along with descriptions. For more information, please refer to the paper below. Our models can be downloaded from the official model repository. 📄 Paper: Small Language Models Learn Enhanced Reasoning Skills from Medical Textbooks Dataset Statistics Table: Statistics of our instruction-tuning datasets“# Examples” denotes the number of training examples for each dataset. †… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/meerkat-instructions.text100K<n<1M11 likes595 downloads1y agoHugging Face02dmisrael /oc20-s2ef-uma-embeddingstabular1M<n<10M0 likes300 downloads1y agoHugging Face03dmis-lab /MedLFQAOriginal dataset introduced by Jeong et al. in OLAPH: Improving Factuality in Biomedical Long-form Question Answering Citation information: @misc{jeong2024olaph, title={OLAPH: Improving Factuality in Biomedical Long-form Question Answering}, author={Minbyul Jeong and Hyeon Hwang and Chanwoong Yoon and Taewhoo Lee and Jaewoo Kang}, year={2024}, eprint={2405.12701}, archivePrefix={arXiv}, primaryClass={cs.CL} } text1K<n<10K17 likes246 downloads2y agoHugging Face04dmis-lab /ChroKnowBench ChroKnowBench ChroKnowBench is a benchmark dataset designed to evaluate the performance of language models on temporal knowledge across multiple domains. The dataset consists of both time-variant and time-invariant knowledge, providing a comprehensive assessment for understanding knowledge evolution and constancy over time. Dataset is introduced by Park et al. in ChroKnowledge: Unveiling Chronological Knowledge of Language Models in Multiple Domains Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/ChroKnowBench.10K<n<100K8 likes142 downloads2y agoHugging Face05dmis-lab /llama-3.1-medprm-reward-training-set Med-PRM-Reward (Version 1.0) 🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-training-set.tabulartext-generation10K<n<100K12 likes135 downloads1y agoHugging Face06dmis-lab /RF-Collection RF_Collection Dataset Description We construct a large-scale dataset called RF-Collection, containing Retrievers' Feedback on oer 410k query rewrites across 12K conversations. Dataset Files The dataset is organized into several CSV files, each corresponding to different retrieval and datasets: TopiOCQA_train_bm25.csv: Contains the retrieval results using the BM25 on the TopiOCQA dataset. TopiOCQA_train_ance.csv: Contains the retrieval results using the ANCE on… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/RF-Collection.1 likes117 downloads1y agoHugging Face07dmis-lab /llama-3.1-medprm-reward-test-set🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its scalability is not limited to… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-test-set.text-generation2 likes77 downloads1y agoHugging Face08dmis-lab /ToxReason ToxReason 🚀 Accepted at ACL 2026 Findings ToxReason is a benchmark dataset for mechanistic chemical toxicity reasoning based on Adverse Outcome Pathways (AOPs). The dataset is designed to evaluate whether large language models can generate biologically interpretable toxicity reasoning trajectories that connect molecular structures, Molecular Initiating Events (MIEs), pathway perturbations, and organ-level adverse outcomes. Dataset Overview ToxReason consists of… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/ToxReason.texttext-generation10K<n<100K0 likes68 downloads4mo agoHugging Face09dmisrael /planned_diffusion_reasoningtext100K<n<1M0 likes56 downloads9mo agoHugging Face10dmis-lab /ETHICGithub : https://github.com/dmis-lab/ETHICArxiv : https://arxiv.org/abs/2410.16848 text1K<n<10K7 likes50 downloads2y agoHugging Face11dmis-lab /llama-3.1-medprm-reward-raw-training-settabular10K<n<100K0 likes31 downloads1y agoHugging Face12dmisrael /pdv2-datatext1M<n<10M0 likes29 downloads9mo agoHugging Face13dmisrael /oc20-s2ef-vae-embeddings-devtabularn<1K0 likes28 downloads1y agoHugging Face14dmis-lab /TemporalHead [ACL 2025] Does Time Have Its Place? Temporal Heads: Where Language Models Recall Time-specific Information This repository contains two separate subsets of data (configs): Temporal: JSON files in Temporal that include temporal knowledge. Invariant: JSON files in Invariant that describe time-invariant knowledge based on LRE. Each subset has its own schema. By defining them as two configs in the YAML header above, Hugging Face’s Dataset Viewer will show “Temporal” and… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/TemporalHead.textquestion-answeringn<1K1 likes25 downloads1y agoHugging Face15dgambettaphd /D_mis72_run0_gen8_WXS_doc1000_synt64_lr1e-04_acm_FRESHtabular10K<n<100K0 likes25 downloads6mo agoHugging Face16dgambettaphd /D_mis72_run0_gen7_WXS_doc1000_synt64_lr1e-04_acm_FRESHtabular10K<n<100K0 likes22 downloads6mo agoHugging Face17dgambettaphd /D_mis72_run0_gen1_WXS_doc1000_synt64_lr1e-04_acm_MPPtabular10K<n<100K0 likes18 downloads6mo agoHugging Face18dmisrael /essential-web-subsample essential-web-subsample A small, category-balanced subsample of Essential-Web v1.0 for studying how generation behavior (e.g. diffusion-LM parallelism) varies with data type. 900 documents = 100 per category across 9 categories, mapped from Essential-Web's document_type_v1.primary.label: category Essential-Web document_type_v1 code Code/Software papers Academic/Research encyclopedic Reference/Encyclopedic/Educational legal Legal/Regulatory literary… See the full description on the dataset page: https://huggingface.co/datasets/dmisrael/essential-web-subsample.tabularn<1K0 likes17 downloads4mo agoHugging Face19dmisrael /oc20-s2ef-vae-embeddings-testtabularn<1K0 likes15 downloads1y agoHugging Face20dmisrael /oc20-s2ef-uma-embeddings-devtabularn<1K0 likes13 downloads1y agoHugging Face21dmisrael /oc20-s2ef-vae-embeddings-test4tabularn<1K0 likes13 downloads1y agoHugging Face22dmisrael /pdv2_testtextn<1K0 likes13 downloads9mo agoHugging Face23dgambettaphd /D_mis_run3_gen8_WXS_doc1000_synt64_lr1e-04_acm_SYNLASTtabular10K<n<100K0 likes13 downloads7mo agoHugging Face24dgambettaphd /D_mis72_run0_gen2_WXS_doc1000_synt64_lr1e-04_acm_MPPtabular10K<n<100K0 likes13 downloads6mo agoHugging Face25dmisrael /pdv2_test_1.2textn<1K0 likes12 downloads9mo agoHugging Face26dgambettaphd /D_mis72_run0_gen4_WXS_doc1000_synt64_lr1e-04_acm_FRESHtabular10K<n<100K0 likes12 downloads6mo agoHugging Face27dgambettaphd /D_mis73_run0_gen3_WXS_doc1000_synt64_lr1e-04_acm_FRESHtabular10K<n<100K0 likes12 downloads6mo agoHugging Face28dgambettaphd /D_mis73_run0_gen4_WXS_doc1000_synt64_lr1e-04_acm_FRESHtabular10K<n<100K0 likes12 downloads6mo agoHugging Face29dmisrael /pdv2_test_1.1textn<1K0 likes11 downloads9mo agoHugging Face30dgambettaphd /D_mis72_run0_gen10_WXS_doc1000_synt64_lr1e-04_acm_FRESHtabular10K<n<100K0 likes11 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.