datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmlu-prox-eval-predictions
MMLU-ProX Multilingual Model Predictions
Raw per-sample model predictions on MMLU-ProX
across 29 languages and 25 open-weight LLMs, produced with
lm-evaluation-harness.
This dataset releases the full prediction logs (not just aggregate scores) so that
item-level responses can be re-analysed — e.g. for Item Response Theory (IRT) modelling
of multilingual benchmarks, error analysis, or per-item difficulty estimation.
Repository structure
mmlu_prox_<lang>/
└──… See the full description on the dataset page: https://huggingface.co/datasets/gililior/mmlu-prox-eval-predictions.clinical-trial-outcomes-predictions
Clinical Trial Outcomes Prediction Dataset
A dataset of 1,366 binary forecasting questions about clinical trial outcomes, automatically generated and labeled using Lightning Rod Labs' Future-as-Label methodology.
Dataset Description
This dataset contains questions about pharmaceutical clinical trials from 2023-2024, paired with verified outcomes (success/failure). Each question asks whether a specific trial will meet its endpoints, receive FDA approval, or complete by a… See the full description on the dataset page: https://huggingface.co/datasets/3rdSon/clinical-trial-outcomes-predictions.symptom-based-disease-prediction-v1
Symptom-Based Disease Prediction Dataset (Confidence-Aware) – v1
📘 Overview
This dataset is an early-stage (Version 1) medical dataset created for symptom-based disease prediction using Large Language Models (LLMs).
Each record presents patient symptoms in an instruction-style prompt and returns multiple possible diseases grouped by confidence levels.The primary goal of this version is to establish structure, consistency, and reasoning format, not final model… See the full description on the dataset page: https://huggingface.co/datasets/Jainam-11/symptom-based-disease-prediction-v1.symptom-based-disease-prediction-v2
Symptom-Based Disease Prediction Dataset (Confidence-Aware) – v2
📘 Overview
Version 2 of the Symptom-Based Disease Prediction Dataset is an improved and more structured medical reasoning dataset designed for Large Language Models (LLMs) and AI-driven healthcare research.
This version introduces:
cleaner formatting,
improved disease grouping,
better confidence separation,
enhanced consistency,
and more fine-tuning-friendly outputs.
Each sample presents symptoms in… See the full description on the dataset page: https://huggingface.co/datasets/Jainam-11/symptom-based-disease-prediction-v2.frames-benchmark-predictions
BENCH-04: N-8 Research Google FRAMES Benchmark Evaluation
Team Designation: N-8 ResearchLead Author: Greg VivianoOrganization: N-8 ResearchEvaluated Dataset: google/frames-benchmark (824 Questions)
Executive Summary
This repository contains the prediction dataset generated by N-8 Research's Deterministic Context Architecture across all 824 multi-step enterprise reasoning questions in Google's official google/frames-benchmark.
Performance Scorecard… See the full description on the dataset page: https://huggingface.co/datasets/gviviano/frames-benchmark-predictions.qwen-math-predictions
qwen-math-predictions
Dataset Description
Qwen MATH training predictions (no confidence tags)
Model: Qwen/Qwen2.5-7B-Instruct
Split: Training set predictions
Dataset Structure
Features
response: Model-generated response
question: Original MATH problem
correct_answer: Ground truth answer
Data Size
Total examples: 6,750
Sample
{
"response": "We start with the given equation:\n\\[ 4^6 = 8^n \\]\n\nFirst, we express both 4 and… See the full description on the dataset page: https://huggingface.co/datasets/huyxdang/qwen-math-predictions.symptom-based-disease-prediction-v2
Symptom-Based Disease Prediction Dataset (Confidence-Aware) – v1
📘 Overview
This dataset is an early-stage (Version 1) medical dataset created for symptom-based disease prediction using Large Language Models (LLMs).
Each record presents patient symptoms in an instruction-style prompt and returns multiple possible diseases grouped by confidence levels.The primary goal of this version is to establish structure, consistency, and reasoning format, not final model… See the full description on the dataset page: https://huggingface.co/datasets/RonalLI/symptom-based-disease-prediction-v2.
