datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmlu-prox-eval-predictions
MMLU-ProX Multilingual Model Predictions
Raw per-sample model predictions on MMLU-ProX
across 29 languages and 25 open-weight LLMs, produced with
lm-evaluation-harness.
This dataset releases the full prediction logs (not just aggregate scores) so that
item-level responses can be re-analysed — e.g. for Item Response Theory (IRT) modelling
of multilingual benchmarks, error analysis, or per-item difficulty estimation.
Repository structure
mmlu_prox_<lang>/
└──… See the full description on the dataset page: https://huggingface.co/datasets/gililior/mmlu-prox-eval-predictions.clinical-trial-outcomes-predictions
Clinical Trial Outcomes Prediction Dataset
A dataset of 1,366 binary forecasting questions about clinical trial outcomes, automatically generated and labeled using Lightning Rod Labs' Future-as-Label methodology.
Dataset Description
This dataset contains questions about pharmaceutical clinical trials from 2023-2024, paired with verified outcomes (success/failure). Each question asks whether a specific trial will meet its endpoints, receive FDA approval, or complete by a… See the full description on the dataset page: https://huggingface.co/datasets/3rdSon/clinical-trial-outcomes-predictions.frames-benchmark-predictions
BENCH-04: N-8 Research Google FRAMES Benchmark Evaluation
Team Designation: N-8 ResearchLead Author: Greg VivianoOrganization: N-8 ResearchEvaluated Dataset: google/frames-benchmark (824 Questions)
Executive Summary
This repository contains the prediction dataset generated by N-8 Research's Deterministic Context Architecture across all 824 multi-step enterprise reasoning questions in Google's official google/frames-benchmark.
Performance Scorecard… See the full description on the dataset page: https://huggingface.co/datasets/gviviano/frames-benchmark-predictions.
