datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sardi-data
SARDI — Evaluation Data
Test splits and prebuilt BM25 indices for Self-Augmenting Retrieval for
Diffusion Language Models (ICML 2026).
Paper · Code · Model
Download
hf download pauljngr/sardi-data --repo-type dataset --local-dir data
Contents
dataset
questions
passages
size
2WikiMultiHopQA
6,253
406,822
308 MB
HotpotQA
3,701
5,239,002
2.7 GB
MuSiQue
2,417
103,035
92 MB
CofCA
900
3,156
6 MB
SynthWorlds-SM
1,200
8,055
16 MB… See the full description on the dataset page: https://huggingface.co/datasets/pauljngr/sardi-data.physiology-mcqa-8kThis dataset is a subset of MedMCQA
