datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medpmc-11m-dataset_jun24_baseline
MedPMC WebDataset
MedPMC is a large-scale medical image-text dataset curated from articles in the PubMed Central (PMC) collection. This release contains approximately 11 million image-text pairs collected from the June 2024 PMC baseline. MedPMC is an ongoing effort, and future releases will continue to expand the dataset with newly published literature, improved annotations, and additional resources.
This dataset is presented in the paper MedPMC: A Systematic Framework for… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/medpmc-11m-dataset_jun24_baseline.objaverse-lvissd2.1_base_traincommonsense-baselineMozzaVID_Base
MozzaVID dataset - Base split
A dataset of synchrotron X-ray tomography scans of mozzarella microstructure, aimed for volumetric model benchmarking and food structure analysis.
[Paper] [Project website]
This version is prepared in the WebDataset format, optimized for streaming. Check our GitHub for details on how to use it. To download raw data instead, visit: [LINK].
Dataset splits
This is a Base split of the dataset containing 4 728 volumes. We also… See the full description on the dataset page: https://huggingface.co/datasets/dtudk/MozzaVID_Base.GPQA_verifications_GenRM-Base_Llama-3.3-70B-Instructbaseline_3_v4_collection_assets
baseline_3 v4 — sim collection assets
2026-05-26 — STATUS: ARCHIVED / REFERENCE-ONLY.
The plan to run sim collection on the A6000 was abandoned this round
due to a glibc incompatibility: the A6000 system glibc is 2.31 (Ubuntu
20.04) but IsaacSim 5.1.0 requires glibc 2.35 (Ubuntu 24.04). All sim
collection (DexYCB + OakInk) was completed on the dev box (RTX 5090).
The contents of this dataset (tarballs of source episodes, USDs, scripts,
partial results log) are kept for reference… See the full description on the dataset page: https://huggingface.co/datasets/UCBProject/baseline_3_v4_collection_assets.web3-knowledge-basecompressed_verifications_lcb128_llama-3.3-70b_genrm_baseMATH128_verifications_Llama-3.3-70B-Instruct_GenRM-Baseacwm-baseline-resultspiper-plus-base-dataset
Piper Plus Base Dataset (6-Language Multilingual)
Piper Plus TTS の事前学習用 6 言語マルチリンガルデータセットです。音素化・スペクトログラム計算済みのため、ダウンロード後すぐに学習を開始できます。
Dataset Summary
Item
Value
Total utterances
497,519
Speakers
571
Languages
6
Phoneme symbols
173
Sample rate
22,050 Hz
Format
WebDataset (tar shards)
Shards
48
Total size
~411 GB
Per-Language Breakdown
Language
Code
ID
Speakers
Utterances
Source
Japanese
ja
0
20
59,694
MOE-Speech… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/piper-plus-base-dataset.worldrenderer-baseline-outputbaselines_videosdxl-turbo-fp-baselineVista4D_baseline
