datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SoundnessBench
SoundnessBench
SoundnessBench is a benchmark of 1,099 machine learning research proposals reconstructed from ICLR submissions and labeled with reviewer soundness sub-scores.
It is intended to evaluate whether LLMs can judge proposal-stage methodological soundness before expensive experimentation.
Data
The dataset file is:
soundnessbench.jsonl
Expected local layout:
data/soundnessbench.jsonl
Evaluation
Reference code and prompts live here:… See the full description on the dataset page: https://huggingface.co/datasets/hosytuyen/SoundnessBench.soundness_data_integrity.jsonsoundness_logic_validation.jsonsoundness_audio_events.json
