CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01junma /CVPR-BiomedSegFMThis repository contains the BiomedSegFM dataset, a crucial resource for the CVPR 2025 Competition: Foundation Models for 3D Biomedical Image Segmentation. Foundation Models for Interactive 3D Biomedical Image Segmentation (Homepage) Foundation Models for Text-guided 3D Biomedical Image Segmentation (Homepage) CVPR 2025 Competition: Foundation Models for 3D Biomedical Image Segmentation Highly recommend watching the webinar recording to learn about the task settings and… See the full description on the dataset page: https://huggingface.co/datasets/junma/CVPR-BiomedSegFM.3dimage-segmentation24 likes35k downloads6mo agoHugging Face02BIOMEDICA /biomedica_webdataset_24Mgated Dataset Card for Dataset Name Arxiv: Arxiv &nbsp;&nbsp;&nbsp;&nbsp;|&nbsp;&nbsp;&nbsp;&nbsp; Website: Biomedica &nbsp;&nbsp;&nbsp;&nbsp;|&nbsp;&nbsp;&nbsp;&nbsp; Training instructions: OpenCLIP &nbsp;&nbsp;&nbsp;&nbsp;|&nbsp;&nbsp;&nbsp;&nbsp; Tutorial: Google Colab BIOMEDICA Dataset is a large-scale, deep-learning-ready biomedical dataset containing over 24M imagecaption pairs and 30M image-references from 6M unique open-source articles. Each… See the full description on the dataset page: https://huggingface.co/datasets/BIOMEDICA/biomedica_webdataset_24M.n>1T40 likes9.2k downloads21d agoHugging Face03Biomedical-TeMU /SPACCC_Tokenizer The Tokenizer for Clinical Cases Written in Spanish Introduction This repository contains the tokenization model trained using the SPACCC_TOKEN corpus (https://github.com/PlanTL-SANIDAD/SPACCC_TOKEN). The model was trained using the 90% of the corpus (900 clinical cases) and tested against the 10% (100 clinical cases). This model is a great resource to tokenize biomedical documents, specially clinical cases written in Spanish. This model was created using the Apache… See the full description on the dataset page: https://huggingface.co/datasets/Biomedical-TeMU/SPACCC_Tokenizer.text10K<n<100K0 likes4.5k downloads5y agoHugging Face04vidore /biomedical_lectures_v2 Vidore Benchmark 2 - MIT Dataset (Multilingual) This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of MIT courses in anatomy (precisely tissue interactions). Dataset Summary The dataset contain queries in the following languages : ["english", "french", "german", "spanish"]. Each query was originaly in "english" (see… See the full description on the dataset page: https://huggingface.co/datasets/vidore/biomedical_lectures_v2.imagedocument-question-answering1K<n<10K0 likes4.1k downloads1y agoHugging Face05Biomedical-TeMU /CodiEsp_corpus Introduction These are the train, development, test and background sets of the CodiEsp corpus. Train and development have gold standard annotations. The unannotated background and test sets are distributed together. All documents are released in the context of the CodiEsp track for CLEF ehealth 2020 (http://temu.bsc.es/codiesp/). The CodiEsp corpus contains manually coded clinical cases. All documents are in Spanish language and CIE10 is the coding terminology (it is the Spanish… See the full description on the dataset page: https://huggingface.co/datasets/Biomedical-TeMU/CodiEsp_corpus.text10K<n<100K0 likes2.1k downloads5y agoHugging Face06almanach /Biomed-Enriched Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content Dataset Authors Rian Touchent, Nathan Godey & Eric de la ClergerieSorbonne Université, INRIA Paris Overview Biomed-Enriched is a PubMed-derived dataset created using a two-stage annotation process. Initially, Llama 3.1 70B Instruct annotated 400K paragraphs for document type, domain, and educational quality. These annotations were then… See the full description on the dataset page: https://huggingface.co/datasets/almanach/Biomed-Enriched.tabulartext-classification100M<n<1B9 likes1.4k downloads2mo agoHugging Face07zou-lab /BioMed-R1-Eval Disentangling Reasoning and Knowledge in Medical Large Language Models This is the evaluation dataset accompanying our paper, comprising 11 publicly available biomedical benchmarks. We disentangle each benchmark question into either medical reasoning or medical knowledge categories. Additionally, we provide a set of adversarial reasoning traces designed to evaluate the robustness of medical reasoning models. For more details, please refer to our GitHub. If you find this work useful… See the full description on the dataset page: https://huggingface.co/datasets/zou-lab/BioMed-R1-Eval.tabularquestion-answering10K<n<100K1 likes1.1k downloads1y agoHugging Face08Biomedical-TeMU /ProfNER_corpus_classificationtext10K<n<100K2 likes1.1k downloads5y agoHugging Face09Biomedical-TeMU /ProfNER_corpus_NER Description Gold standard annotations for profession detection in Spanish COVID-19 tweets The entire corpus contains 10,000 annotated tweets. It has been split into training, validation, and test (60-20-20). The current version contains the training and development set of the shared task with Gold Standard annotations. In addition, it contains the unannotated test, and background sets will be released. For Named Entity Recognition, profession detection, annotations are distributed… See the full description on the dataset page: https://huggingface.co/datasets/Biomedical-TeMU/ProfNER_corpus_NER.text1 likes991 downloads5y agoHugging Face10Novel-BioMedAI /Medical_Segmentation_Decathlon 🏆 Medical Segmentation Decathlon Dataset 📝 Overview The Medical Segmentation Decathlon (MSD) is a comprehensive benchmark dataset for validating algorithms in 3D medical image segmentation. It includes 10 distinct tasks, each with unique challenges like small data sizes, unbalanced labels, varying object scales, multi-class labels, and multimodal imaging. 🔗 Dataset Access 🌐 Website: Medical Decathlon 📂 Google Drive: MSD Google Drive 🧩 Task… See the full description on the dataset page: https://huggingface.co/datasets/Novel-BioMedAI/Medical_Segmentation_Decathlon.2 likes918 downloads1y agoHugging Face11Biomedical-TeMU /SPACCC_Sentence-Splitter The Sentence Splitter (SS) for Clinical Cases Written in Spanish Introduction This repository contains the sentence splitting model trained using the SPACCC_SPLIT corpus (https://github.com/PlanTL-SANIDAD/SPACCC_SPLIT). The model was trained using the 90% of the corpus (900 clinical cases) and tested against the 10% (100 clinical cases). This model is a great resource to split sentences in biomedical documents, specially clinical cases written in Spanish. This model… See the full description on the dataset page: https://huggingface.co/datasets/Biomedical-TeMU/SPACCC_Sentence-Splitter.text10K<n<100K1 likes869 downloads5y agoHugging Face12Alejandro98 /Biomedica2025EvalSet Biomedica 2025 Eval Set Unified test snapshot of the BioMedica 2025 vision–language evaluation suites used in AMInZeroShotOpenEvalAllTasks. Every row is a single image with closed-ended options, the gold answer, and provenance fields that point back to the original dataset. Images are stored as original JPEG/PNG bytes (or JPEG-encoded arrays) inside parquet so the Hugging Face dataset viewer is enabled (~1959 MB download, 89941 examples). Suites config… See the full description on the dataset page: https://huggingface.co/datasets/Alejandro98/Biomedica2025EvalSet.imageimage-classification100K<n<1M0 likes583 downloads1mo agoHugging Face13biomedical-translator /pubmed2024_sentence_embeddingstext100M<n<1B0 likes412 downloads2y agoHugging Face14li-lab /HLE-BioMedX HLE-BioMedX — Multilingual HLE Biology/Medicine A multilingual version of the Biology/Medicine subset of Humanity's Last Exam (HLE), released as one subset per language. Source benchmark: Humanity's Last Exam, dataset cais/hle. Subsets Group Languages How the target-language text was produced Source en Original English questions and answers. Machine-translated and expert-verified / revised zh, ja, ko, fr, th Machine translation reviewed by a human… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/HLE-BioMedX.textquestion-answering1K<n<10K0 likes382 downloads25d agoHugging Face15open-llm-leaderboard-old /details_lighteternal__Llama3-merge-biomed-8b0 likes380 downloads2y agoHugging Face16TahaKoleilat /BiomedCoOp Biomedical Few-shot Image Classification for Vision-Language Models Overview Recent advancements in vision-language models (VLMs), such as CLIP, have demonstrated substantial success in self-supervised representation learning for vision tasks. However, effectively adapting VLMs to downstream applications remains challenging, as their accuracy often depends on time-intensive and expertise-demanding prompt engineering, while full model fine-tuning is… See the full description on the dataset page: https://huggingface.co/datasets/TahaKoleilat/BiomedCoOp.zero-shot-image-classification1 likes364 downloads4mo agoHugging Face17rntc /biomed-fr Dataset Card for "biomed-fr" More Information needed text10M<n<100M0 likes342 downloads4y agoHugging Face18anthonyyazdaniml /gliner-biomed-pre-training GLiNER-BioMed pre-training dataset This dataset, used for the pre-training stage of the GLiNER-BioMed models, was introduced in the paper GLiNER-BioMed: a suite of efficient models for open biomedical named entity recognition. Citation If you use the GLiNER-BioMed models or datasets in your work, please cite: @article{yazdani2026gliner, author = {Yazdani, Anthony and Stepanov, Ihor and Teodoro, Douglas}, title = {{GLiNER-BioMed}: a suite of efficient… See the full description on the dataset page: https://huggingface.co/datasets/anthonyyazdaniml/gliner-biomed-pre-training.text10K<n<100K0 likes309 downloads2mo agoHugging Face19vidore /biomedical_lectures_eng_v2 Vidore Benchmark 2 - MIT Dataset This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of MIT courses in anatomy (precisely tissue interactions). Dataset Summary Each query is in english. This dataset provides a focused benchmark for visual retrieval tasks related to MIT biology courses. It includes a curated set of documents, queries, relevance judgments (qrels), and page images.… See the full description on the dataset page: https://huggingface.co/datasets/vidore/biomedical_lectures_eng_v2.imagedocument-question-answering1K<n<10K0 likes306 downloads1y agoHugging Face20disi-unibo-nlp /Pile-NER-biomed-IOBtext10K<n<100K0 likes277 downloads1y agoHugging Face21DataTecnica /RoP_biomedical RoP — Biomedical Reference of Parameters One CDE set to rule them all... 📋 Interested in help using RoP? → Submit Interest Form ← Start here Executive Summary RoP decreases activation energy in biomedical research by accelerating progress past the biggest bottlenecks ... data wrangling, interoperability and AI-readiness. Right now, multi-cohort studies spend months (often years) manually mapping variables—PPMI calls it "MoCA Total Score," NACC calls it… See the full description on the dataset page: https://huggingface.co/datasets/DataTecnica/RoP_biomedical.tabularfeature-extraction1M<n<10M3 likes260 downloads2mo agoHugging Face22simpleG2023 /chinese-biomedicine-and-genomics-open-intelligence 🔬 Chinese Biomedicine, Cell Therapy & Genomics Open Intelligence Dataset Curated open intelligence dataset providing English briefs, clinical trial benchmarks, verified abstracts, and DOIs of frontier Chinese research in Cellular Therapeutics, Gene Editing, ADCs, and NMPA Clinical Approvals. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-biomedicine-and-genomics-open-intelligence.tabulartext-retrieval1K<n<10K0 likes255 downloads22h agoHugging Face23rntc /multi_biomed_french_edu3_20250710text10K<n<100K0 likes252 downloads1y agoHugging Face24anthonyyazdaniml /gliner-biomed-post-training GLiNER-BioMed post-training dataset This dataset, used for the post-training stage of the GLiNER-BioMed models, was introduced in the paper GLiNER-BioMed: a suite of efficient models for open biomedical named entity recognition. Citation If you use the GLiNER-BioMed models or datasets in your work, please cite: @article{yazdani2026gliner, author = {Yazdani, Anthony and Stepanov, Ihor and Teodoro, Douglas}, title = {{GLiNER-BioMed}: a suite of efficient… See the full description on the dataset page: https://huggingface.co/datasets/anthonyyazdaniml/gliner-biomed-post-training.text10K<n<100K0 likes250 downloads2mo agoHugging Face25anthonyyazdaniml /gliner-biomed-curated-corpus GLiNER-BioMed curated corpus Unlabeled corpus introduced in the paper GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition. Citation If you use the GLiNER-BioMed models or datasets in your work, please cite: @misc{yazdani2025glinerbiomedsuiteefficientmodels, title={GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition}, author={Anthony Yazdani and Ihor Stepanov and Douglas Teodoro}… See the full description on the dataset page: https://huggingface.co/datasets/anthonyyazdaniml/gliner-biomed-curated-corpus.text100K<n<1M0 likes230 downloads1y agoHugging Face26rntc /multi_biomed_french_prefixed_20250710text10K<n<100K0 likes229 downloads1y agoHugging Face27anthonyyazdaniml /gliner-biomed-balanced-curated-corpus GLiNER-BioMed balanced curated corpus Balanced, unlabeled corpus introduced in the paper GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition. Citation If you use the GLiNER-BioMed models or datasets in your work, please cite: @misc{yazdani2025glinerbiomedsuiteefficientmodels, title={GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition}, author={Anthony Yazdani and Ihor Stepanov and Douglas… See the full description on the dataset page: https://huggingface.co/datasets/anthonyyazdaniml/gliner-biomed-balanced-curated-corpus.text100K<n<1M0 likes204 downloads1y agoHugging Face28BIOMEDICA /BioMedFlickr BioMedFlickr BioMedFlickr is a biomedical image–caption retrieval benchmark built from public Flickr pathology / microscopy albums. Each example is a single image paired with a cleaned English caption. Images are stored as original JPEGs inside parquet (~865 MB download). This dataset is the retrieval benchmark used in BIOMEDICA (Lozano et al.), published at CVPR 2025: https://cvpr.thecvf.com/virtual/2025/poster/33761 Load from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/BIOMEDICA/BioMedFlickr.imageimage-to-text1K<n<10K2 likes203 downloads21d agoHugging Face29NIH-CARD /BiomedSQLBiomedSQL GitHub Paper Dataset Summary BiomedSQL is a text-to-SQL benchmark designed to evaluate Large Language Models (LLMs) on scientific tabular reasoning tasks. It consists of curated question-SQL query-answer triples covering a variety of biomedical and SQL reasoning types. The benchmark challenges models to apply implicit scientific criteria rather than simply translating syntax. Repository Organization benchmark_data: contains the question-SQL query-answer triples. Please note that you… See the full description on the dataset page: https://huggingface.co/datasets/NIH-CARD/BiomedSQL.texttable-question-answering10K<n<100K4 likes202 downloads6mo agoHugging Face30knowledgator /biomed_NER Biomed NER This dataset consists of 4,840 manually annotated text records drawn from PubMed abstracts, drug descriptions from the FDA, and patent abstracts. All entities are continuous, and there are no nested entities. Dataset composition The dataset contains 4,840 annotated text records distributed across three sources: Source Approx. records Purpose PubMed abstracts ~4,300 Core biomedical content FDA drug descriptions ~430 Pharmaceutical text with dense… See the full description on the dataset page: https://huggingface.co/datasets/knowledgator/biomed_NER.texttoken-classification1K<n<10K12 likes198 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.