datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmu_legacysurvey_dr10_south_21
mmu_legacysurvey_dr10_south_21 HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_legacysurvey_dr10_south_21.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_legacysurvey_dr10_south_21.arc-aphasia-bids
Aphasia Recovery Cohort (ARC)
Multimodal neuroimaging dataset for stroke-induced aphasia research.
Dataset Summary
The Aphasia Recovery Cohort (ARC) is a large-scale, longitudinal neuroimaging dataset containing multimodal MRI scans from 230 chronic stroke patients with aphasia. This HuggingFace-hosted version provides direct Python access to the BIDS-formatted data with embedded NIfTI files.
Metric
Count
Subjects
230
Sessions
902
T1-weighted scans
444… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/arc-aphasia-bids.isles24-stroke
ISLES'24 Stroke Training Dataset
Multi-center longitudinal multimodal acute ischemic stroke training dataset from the ISLES'24 Challenge.
Overview
149 acute ischemic stroke training cases with:
Admission imaging (ses-01): Non-contrast CT, CT angiography, 4D CT perfusion
Follow-up imaging (ses-02): Post-treatment MRI (DWI, ADC)
Clinical data: Demographics, patient history, admission NIHSS, 3-month mRS outcomes
Annotations: Infarct masks, large vessel occlusion masks… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/isles24-stroke.G1-Pretrain-Data-Repommu_manga
mmu_manga HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_manga.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be installed via… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_manga.AI4FA-Diabimmune
Food Allergy Microbiome Dataset (Experimental)
Dataset Summary
This is an experimental microbiome dataset designed for exploratory research in food allergy classification. The dataset contains multiple data modalities (DNA embeddings, microbiome embeddings, raw DNA sequences) collected longitudinally at several timepoints.
Warning: This dataset is experimental. Its structure is frozen for ongoing research, and it is not ready for benchmarking.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/AI4FA-Diabimmune.opendata
OpenData Consortium
Three open datasets exported from the OpenData Consortium data platform.
Config
Description
Primary format
companies
~103M global companies with firmographic attributes
Parquet
locations
~273M business locations with address and geo data
Parquet
people
~101M business contacts linked to companies
Parquet
Usage
from datasets import load_dataset
companies = load_dataset("OpenDataFoundation/opendata", "companies")
locations =… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/opendata.mmu_apogee_dr17
mmu_apogee_dr17 HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_apogee_dr17.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_apogee_dr17.mmu_hsc_pdr3_wide_21
mmu_hsc_pdr3_wide_21 HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_hsc_pdr3_wide_21.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_hsc_pdr3_wide_21.AI4FA-Goldberg
Food Allergy Microbiome Dataset (Experimental)
Dataset Summary
This is an experimental microbiome dataset designed for exploratory research in food allergy classification. The dataset contains multiple data modalities (DNA embeddings, microbiome embeddings, raw DNA sequences) collected longitudinally at several timepoints.
Warning: This dataset is experimental. Its structure is frozen for ongoing research, and it is not ready for benchmarking.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/AI4FA-Goldberg.CodVa-2-Tokenizer-v2gut-microbiome-allergy-data
Dataset Card for Gut Microbiome–Food Allergy Prediction Datasets
Dataset Summary
This repository contains multiple human gut microbiome datasets curated for predicting food allergy development.
Each dataset corresponds to a distinct cohort with longitudinal microbiome sampling, providing both metadata and derived embeddings suitable for machine learning.
The datasets are designed to support binary classification of subjects into healthy vs allergic categories, enabling… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/gut-microbiome-allergy-data.CodVa-2-Tokenizer-v3aomic-piop1AI4FA-Tanaka
Food Allergy Microbiome Dataset (Experimental)
Dataset Summary
This is an experimental microbiome dataset designed for exploratory research in food allergy classification. The dataset contains multiple data modalities (DNA embeddings, microbiome embeddings, raw DNA sequences) collected longitudinally at several timepoints.
Warning: This dataset is experimental. Its structure is frozen for ongoing research, and it is not ready for benchmarking.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/AI4FA-Tanaka.requestmmu_euclid_q1AI4FA-Gadir
Food Allergy Microbiome Dataset (Experimental)
Dataset Summary
This is an experimental microbiome dataset designed for exploratory research in food allergy classification. The dataset contains multiple data modalities (DNA embeddings, microbiome embeddings, raw DNA sequences) collected longitudinally at several timepoints.
Warning: This dataset is experimental. Its structure is frozen for ongoing research, and it is not ready for benchmarking.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/AI4FA-Gadir.jain-clinical-antibody-polyreactivity
Jain Clinical Antibody Polyreactivity Dataset (Novo Nordisk Parity Benchmark)
Dataset Summary
This dataset contains 86 clinical-stage antibody heavy chain variable domain (VH) sequences with binary polyreactivity labels, preprocessed to reproduce the benchmark results from Sakhnini et al. 2025 (Novo Nordisk & University of Cambridge). The original dataset was published by Jain et al. 2017 and contains biophysical measurements for 137 FDA-approved or late-stage… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/jain-clinical-antibody-polyreactivity.neb-mechanism-elucidation
neb-mechanism-elucidation
Development host for the agent-visible input archive of the Terminal-Bench Science task
neb-mechanism-elucidation.
This repository is a file archive, not a tabular dataset. The Hugging Face dataset
viewer is disabled on purpose.
Contents
Only neb-mechanism-elucidation/input/ is published here. That prefix contains the
sanitized VASP evidence the agent may see. It does not contain the hidden
reference, verifier files, or answer-bearing… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/neb-mechanism-elucidation.jain-developability-cleandataset-quest-indexawesome-food-allergy-datasets
Awesome Food Allergy Datasets
A curated collection of datasets, databases, and computational resources for food allergy research, allergen identification, drug development, and clinical applications.
🧬 Dataset Description
Dataset Summary
Food allergy affects over 220 million people worldwide. This repository serves as the first comprehensive, open collection of AI-ready datasets for food allergy research—spanning clinical trials, immunotherapy, genomics… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/awesome-food-allergy-datasets.m-boltz-submissionsshehata-antibody-psr
Shehata Antibody PSR Dataset (Novo Nordisk Preprocessing)
Dataset Summary
This dataset contains 398 human antibody heavy chain variable domain (VH) sequences with PSR (Poly-Specificity Reagent) measurements, preprocessed according to the methodology described in Sakhnini et al. 2025 (Novo Nordisk & University of Cambridge). The dataset was originally published by Shehata et al. 2019 and contains human B cell-derived antibodies studying the relationship between affinity… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/shehata-antibody-psr.boughter-antibody-polyreactivity
Boughter Antibody Polyreactivity Dataset (Novo Nordisk Preprocessing)
Dataset Summary
This dataset contains 914 antibody heavy chain variable domain (VH) sequences with binary polyreactivity labels, preprocessed according to the methodology described in Sakhnini et al. 2025 (Novo Nordisk & University of Cambridge). The dataset was originally published by Boughter et al. 2020 and contains mouse antibodies with ELISA-based polyreactivity measurements against a panel of 4–7… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/boughter-antibody-polyreactivity.antibody-dev-jain-2017goa_reasoning
Dataset Description
This dataset integrates Gene Ontology (GO) annotations (filtered to experimental evidence, IDA) with bibliographic data from Europe PMC.Each record links a biological entity (gene/protein) to a GO term and the supporting scientific publication (Title and Abstract).
Columns
DB, DB_Object_ID, DB_Object_Symbol — Source database and gene/protein identifiers
Qualifier — Relationship qualifier (e.g., NOT, contributes_to)
GO_ID, GO_Name — GO term and its… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/goa_reasoning.harvey-nanobody-polyreactivity
Harvey Nanobody Polyreactivity Dataset (Novo Nordisk Preprocessing)
Dataset Summary
This dataset contains 141,021 nanobody (VHH) sequences with binary polyreactivity labels, preprocessed according to the methodology described in Sakhnini et al. 2025 (Novo Nordisk & University of Cambridge). The dataset was originally published by Harvey et al. 2022 and contains synthetic nanobodies assessed by PSR (Poly-Specificity Reagent) assay via FACS sorting and deep sequencing.
This… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/harvey-nanobody-polyreactivity.nova1-pretrain-20T
