datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
STXBP1_PubMed_Central_Multimodal_Dataset
STXBP1 PubMed Central Multimodal Dataset v2 (12-13-2025)
A comprehensive multimodal dataset for training vision-language models on biomedical scientific literature, with focus on STXBP1-related neurological research.
🆕 Version 2 Updates (December 2025)
497,360 training examples (up from ~31K)
170,591 matched figure-image pairs (99.7% match rate)
Full captions preserved (no truncation)
Multiple training formats for different use cases
Validated response lengths for… See the full description on the dataset page: https://huggingface.co/datasets/SkyWhal3/STXBP1_PubMed_Central_Multimodal_Dataset.stxbp1-pubmed-central-fulltext
source_datasets:
- PubMed Central
STXBP1 PubMed Central Full-Text Dataset v2
A comprehensive collection of 31,786 full-text scientific articles from PubMed Central related to STXBP1, synaptic function, and neurological research.
🆕 Version 2 Updates (December 2025)
Complete re-extraction with improved HTML parsing
Full main text with proper section headers
Enhanced metadata extraction
99.7% figure-image matching (see companion multimodal dataset)… See the full description on the dataset page: https://huggingface.co/datasets/SkyWhal3/stxbp1-pubmed-central-fulltext.the-pile-pubmed-central-refined-by-data-juicer
The Pile -- PubMed Central (refined by Data-Juicer)
A refined version of PubMed Central dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 83G).
Dataset Information
Number of samples: 2,694,860 (Keep ~86.96% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-pubmed-central-refined-by-data-juicer.
