bids
Datasets
All datasets matching “bids”arc-aphasia-bids
Aphasia Recovery Cohort (ARC)
Multimodal neuroimaging dataset for stroke-induced aphasia research.
Dataset Summary
The Aphasia Recovery Cohort (ARC) is a large-scale, longitudinal neuroimaging dataset containing multimodal MRI scans from 230 chronic stroke patients with aphasia. This HuggingFace-hosted version provides direct Python access to the BIDS-formatted data with embedded NIfTI files.
Metric
Count
Subjects
230
Sessions
902
T1-weighted scans
444… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/arc-aphasia-bids.medpmc-11m-dataset_jun24_baseline
MedPMC WebDataset
MedPMC is a large-scale medical image-text dataset curated from articles in the PubMed Central (PMC) collection. This release contains approximately 11 million image-text pairs collected from the June 2024 PMC baseline. MedPMC is an ongoing effort, and future releases will continue to expand the dataset with newly published literature, improved annotations, and additional resources.
This dataset is presented in the paper MedPMC: A Systematic Framework for… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/medpmc-11m-dataset_jun24_baseline.M3LLM-data
M3LLM-PMC Training Data
This dataset contains the training data for M3LLM (Medical Multimodal Large Language Model), comprising ~238K high-quality synthetic medical instruction-following samples.
Dataset Description
The data is generated from PubMed Central (PMC) medical literature through a comprehensive 5-stage synthetic data pipeline, covering six diverse medical visual question answering tasks.
Dataset Statistics
File
Samples
Task Type… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/M3LLM-data.hbn-multimodal-bids-datasetus-government-bids-sample
US Government Bids — Historical Sample
10,000 historical US government procurement bids with full titles, descriptions, and structured metadata. All bids in this dataset have due dates at least 30 days in the past — this is a historical corpus for research and ML training, not a live opportunity feed.
Source: ProcureTap — aggregated from 290+ federal, state, and local procurement portals (SAM.gov, Grants.gov, state portals, Bonfire, PlanetBids, Jaggaer, BidNet Direct, DemandStar… See the full description on the dataset page: https://huggingface.co/datasets/ProcureTap/us-government-bids-sample.medpmc-screening-dataset
MedPMC Initial Screening Training/Test Datasets
Overview
This dataset contains the annotations used for the initial screening stage of the MedPMC framework, which aims to identify clinically relevant medical images from biomedical literature.
The training and validation sets are automatically curated using GPT-4o. The test set is manually annotated.
For details on dataset construction, annotation guidelines, and the overall MedPMC pipeline, please refer to our… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/medpmc-screening-dataset.
