datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SDO
🌞 SDO ML-Ready Dataset: AIA and HMI Level-1.5
Overview
This dataset provides machine learning (ML)-ready solar data curated from NASA’s Solar Dynamics Observatory (SDO), covering observations from May 13, 2010, to July 31, 2024. It includes Level-1.5 processed data from:
Atmospheric Imaging Assembly (AIA):
Helioseismic and Magnetic Imager (HMI):
The dataset is designed to facilitate large-scale ML applications in heliophysics, such as solar activity forecasting… See the full description on the dataset page: https://huggingface.co/datasets/harshinde/SDO.core-sdo
ML-Ready Multi-Modal Image Dataset from SDO
Overview
This dataset provides machine learning (ML)-ready solar data curated from NASA’s Solar Dynamics Observatory (SDO), covering observations from May 13, 2010, to Dec 31, 2024. It includes Level-1.5 processed data from: Atmospheric Imaging Assembly (AIA)
and Helioseismic and Magnetic Imager (HMI).
The dataset is designed to facilitate large-scale learning applications in heliophysics, such as space weather forecasting… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/core-sdo.sdoml-lite
SDOML-lite
SDOML-lite is a lightweight alternative to the SDOML dataset specifically designed for machine learning applications in solar physics, providing continuous full-disk images of the Sun with magnetic field and extreme ultraviolet data in several wavelengths. The data source is the Solar Dynamics Observatory (SDO) space telescope, a NASA mission that has been in operation since 2010.
NASA’s SDO mission has generated over 20 petabytes of high-resolution solar imagery… See the full description on the dataset page: https://huggingface.co/datasets/oxai4science/sdoml-lite.sdocx-compatibility
SDOCX Compatibility Corpus
This dataset contains paired Samsung Notes .sdocx documents and reference PDF exports for parser, renderer, and visual-regression testing. Each pair has a numeric ID and descriptive name recorded below.
The .sdocx file is the source fixture. The matching PDF is the expected visible result exported from Samsung Notes; it is a visual reference, not a byte-for-byte rendering requirement.
Files
ID
Source
Reference
Coverage… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/sdocx-compatibility.sdo_mobilesd-orgasmic-c1
epicff series: Based on epicphotogasm, using dreambooth to finetuning. Nice outputs, but a little stiff.
epicmq : I don't remember.
htc: Based on SD-1.5 pruned (my mistake), using dreambooth to finetuning. Very creative model, but very hard to create good images by itself.
merged_ft: Based on a mix of SD-1.5 full (7GB) with 70% on epicphotogasm for structure, using Novel AI finetuning. Good mix between flexibility and polish.
SDOH-NLI
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/SDOH-NLI.sd_outbayc
Dataset Card for "bayc"
More Information needed
longCOVID_CaseReports_SDOHThis dataset contains text for over 7,000 case report sections in academic papers related to Post COVID-19 Condition (PCC), also known as Long COVID. It is comprised of three columns: 'Token,' 'entity_type,' and 'case_report_id.' There are over 3.4 million observations in total.
Entity types are presented in CoNLL format and relate to sociodemographic, behavioral and selected biomedical dimensions relevant to PCC. There are 27 core entity types (53 in CoNLL): ‘Access_To_Care,’ ‘Age,’… See the full description on the dataset page: https://huggingface.co/datasets/PCC-SDOH-NLP/longCOVID_CaseReports_SDOH.narrativeshield-sdoh-medqa
NarrativeShield: SDoH Persona-Varied Clinical QA Dataset
Paper: NarrativeShield: Adversarial Intake Against Narrative Anchoring for Equitable Clinical Diagnosis
Dataset Summary
NarrativeShield is a health equity benchmark for clinical NLP. It presents 1,000 USMLE-style clinical questions, each instantiated across three sociolinguistic persona voices grounded in Social Determinants of Health (SDoH) literature — yielding 3,000 total clinical encounters.
The… See the full description on the dataset page: https://huggingface.co/datasets/Prabhjotschugh/narrativeshield-sdoh-medqa.sdoh-nliSDOH-NLISDOH-NLI is a natural language inference dataset containing ~30k premise-hypothesis pairs with binary entailment labels in the domain of social and behavioral determinants of health.
@misc{lelkes2023sdohnli,
title={SDOH-NLI: a Dataset for Inferring Social Determinants of Health from Clinical Notes},
author={Adam D. Lelkes and Eric Loreaux and Tal Schuster and Ming-Jun Chen and Alvin Rajkomar},
year={2023},
eprint={2310.18431},
archivePrefix={arXiv}… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/SDOH-NLI.sd_ocr_rl_embeddingsthe_amazing_spidermanSDOcatalan-storiesSDOH_CTAsop_datasetsolar-sdoThis repo contains the dataset used to train models from https://github.com/SLAMPAI/generative-models-for-highres-solar-images/ (paper: https://arxiv.org/abs/2304.07169).
The dataset is based on the SDO dataset pre-processed following the paper https://arxiv.org/abs/2304.07169/ (see Section 3).
It contains 38676 images (AIA, channel 193Å) of size 1024x1024. See the paper for more details.
sd-outputsbatmansmu_counsel_smplSDO_AIA_2024_94_193_131_2000_samples
Frame sequences of patches of the Sun's surface preceding flares with an assigned GOES class. SDO AIA @ 94, 193, and 131 Å
torchvision-Sdomainnetsdoprofaerial-sdo-dataset
