datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmlu-prox-eval-predictions
MMLU-ProX Multilingual Model Predictions
Raw per-sample model predictions on MMLU-ProX
across 29 languages and 25 open-weight LLMs, produced with
lm-evaluation-harness.
This dataset releases the full prediction logs (not just aggregate scores) so that
item-level responses can be re-analysed — e.g. for Item Response Theory (IRT) modelling
of multilingual benchmarks, error analysis, or per-item difficulty estimation.
Repository structure
mmlu_prox_<lang>/
└──… See the full description on the dataset page: https://huggingface.co/datasets/gililior/mmlu-prox-eval-predictions.polymarket-predictions
THE ORACLE — Polymarket predictions
Live predictions for Polymarket markets, produced by THE ORACLE — an autonomous
agent funded by $ORACLE pump.fun creator fees. Each row is a baseline-model
forecast over live orderbook signals (momentum, microstructure, liquidity).
predictions.json / predictions.csv — 100 markets, refreshed each agent cycle.
Columns: question, category, market_prob, oracle_prob, edge, confidence, signal,
model, backtest_acc, auc, modelability, volume… See the full description on the dataset page: https://huggingface.co/datasets/THEORACLEEEE/polymarket-predictions.phantom-wiki-v0-5-0-predictions
Dataset Card for Dataset Name
Predictions from https://huggingface.co/datasets/mlcore/phantom-wiki-v050
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/kilian-group/phantom-wiki-v0-5-0-predictions.ner-eval-predictionstrajectory-prediction-waymo
Waymo Trajectory Prediction
Dataset Description
This dataset contains preprocessed trajectory prediction samples for autonomous driving research,
formatted for use with DiscoBench's TrajectoryPrediction task.
Original Dataset: Waymo Open Motion Dataset
Number of Samples: 850
Format: Pickle files with numpy arrays
Task: Multi-modal trajectory prediction
Dataset Structure
Each sample is a pickle file containing:
obj_trajs (32, 21, 2): Past trajectories of… See the full description on the dataset page: https://huggingface.co/datasets/saeedrmd/trajectory-prediction-waymo.employee-burnout-turnover-prediction-800k
Synthetic Employee Dataset
800,000+ employee records with real-world distributions for burnout prediction, turnover analysis, and HR analytics
What You Get
This isn't just another CSV dump. You're looking at 800K+ carefully engineered employee profilesthat mirror actual workforce dynamics, complete with performance metrics, burnout indicators, skill matrices, and behavioral personas. Think of it as a production-ready HR database that never existed but feels like it… See the full description on the dataset page: https://huggingface.co/datasets/BrotherTony/employee-burnout-turnover-prediction-800k.python-execution-trace-output-prediction-cmskdimp
Python Execution Trace & Output Prediction
Self-contained Python programs paired with a concrete function call and the exact runtime output. Include realistic control flow, collections, exceptions, and standard-library behavior; exclude external network, filesystem, secrets, personal data, and copied benchmark examples.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Accepted items: 1000… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/python-execution-trace-output-prediction-cmskdimp.review_helpfulness_prediction
Dataset Card for Review Helpfulness Prediction (RHP) Dataset
Dataset Summary
The success of e-commerce services is largely dependent on helpful reviews that aid customers in making informed purchasing decisions. However, some reviews may be spammy or biased, making it challenging to identify which ones are helpful. Current methods for identifying helpful reviews only focus on the review text, ignoring the importance of who posted the review and when it was posted.… See the full description on the dataset page: https://huggingface.co/datasets/tafseer-nayeem/review_helpfulness_prediction.symptom-based-disease-prediction-v1
Symptom-Based Disease Prediction Dataset (Confidence-Aware) – v1
📘 Overview
This dataset is an early-stage (Version 1) medical dataset created for symptom-based disease prediction using Large Language Models (LLMs).
Each record presents patient symptoms in an instruction-style prompt and returns multiple possible diseases grouped by confidence levels.The primary goal of this version is to establish structure, consistency, and reasoning format, not final model… See the full description on the dataset page: https://huggingface.co/datasets/Jainam-11/symptom-based-disease-prediction-v1.lab2_word_predictiontext-to-sql-eval-predictions
What the text-to-SQL models actually generated
Every prediction behind the numbers in
qwen3-8b-text2sql-qlora: the 453 test
questions of the enterprise text-to-SQL benchmark,
each answered by four configurations of the same model, each answer executed against the reference
PostgreSQL database and scored by comparing result sets. 1,812 rows.
I published this because the headline table (10.82 % → 50.99 % → 52.10 %) is the least interesting part of
that project. The interesting… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-eval-predictions.HealthyLife-Insurance-Charge-Prediction-v2bbnaija2026-predictionsmissing_triple_prediction-small
Given a text and semantic graph, predict the missing triple in semantic graph.
NOTE: if the semantic graph is completed! the output should be "Correct"!
symptom-based-disease-prediction-v2
Symptom-Based Disease Prediction Dataset (Confidence-Aware) – v2
📘 Overview
Version 2 of the Symptom-Based Disease Prediction Dataset is an improved and more structured medical reasoning dataset designed for Large Language Models (LLMs) and AI-driven healthcare research.
This version introduces:
cleaner formatting,
improved disease grouping,
better confidence separation,
enhanced consistency,
and more fine-tuning-friendly outputs.
Each sample presents symptoms in… See the full description on the dataset page: https://huggingface.co/datasets/Jainam-11/symptom-based-disease-prediction-v2.employee-burnout-turnover-prediction-800k
Synthetic Employee Dataset
800,000+ employee records with real-world distributions for burnout prediction, turnover analysis, and HR analytics
What You Get
This isn't just another CSV dump. You're looking at 800K+ carefully engineered employee profilesthat mirror actual workforce dynamics, complete with performance metrics, burnout indicators, skill matrices, and behavioral personas. Think of it as a production-ready HR database that never existed but feels like it… See the full description on the dataset page: https://huggingface.co/datasets/Mystic777/employee-burnout-turnover-prediction-800k.repro-ski-rental-with-distributional-predictions-of-unknown-quality-traces
Agent traces
Agent sessions published from a Trackio Logbook.
Tumbuka_Continuous_Next-Token_Prediction_Datasetarc-bad-prediction
ARC Bad Prediction
Visualization of bad predictions: ARC-AGI training, ARC-AGI evaluation.
My goal during the ARC-AGI contests has been to make a stepwise refinement algorithm, that can improve on earlier predictions.
This repo is intended for stepwise refinement algorithms. This dataset contains incorrect predictions that are somewhat close to the target.
I have manually inspected these predictions and removed the worst predictions. However there may still be more bad predictions.… See the full description on the dataset page: https://huggingface.co/datasets/neoneye/arc-bad-prediction.Bitcoin-Long-Term-Trend-and-Price-Prediction-Datasettags:
financial-forecasting
time-series
instruction-tuning
bitcoin
finance
license: apache-2.0
Bitcoin Long-Term Trend and Price Prediction Dataset
This dataset is designed for fine-tuning language models on a long-term Bitcoin forecasting task. The goal is to predict both the overall price trend (up, down, or no change) and the specific daily closing prices for the next 10 days, based on a comprehensive 60-day historical context.
The dataset provides a rich, multi-factor view for each… See the full description on the dataset page: https://huggingface.co/datasets/tahamajs/Bitcoin-Long-Term-Trend-and-Price-Prediction-Dataset.lab_2_word_predictionframes-benchmark-predictions
BENCH-04: N-8 Research Google FRAMES Benchmark Evaluation
Team Designation: N-8 ResearchLead Author: Greg VivianoOrganization: N-8 ResearchEvaluated Dataset: google/frames-benchmark (824 Questions)
Executive Summary
This repository contains the prediction dataset generated by N-8 Research's Deterministic Context Architecture across all 824 multi-step enterprise reasoning questions in Google's official google/frames-benchmark.
Performance Scorecard… See the full description on the dataset page: https://huggingface.co/datasets/gviviano/frames-benchmark-predictions.Sequence-of-action-prediction-mind2webemployee-burnout-turnover-prediction-800k
Synthetic Employee Dataset
800,000+ employee records with real-world distributions for burnout prediction, turnover analysis, and HR analytics
What You Get
This isn't just another CSV dump. You're looking at 800K+ carefully engineered employee profilesthat mirror actual workforce dynamics, complete with performance metrics, burnout indicators, skill matrices, and behavioral personas. Think of it as a production-ready HR database that never existed but feels like it… See the full description on the dataset page: https://huggingface.co/datasets/Umer112233/employee-burnout-turnover-prediction-800k.PubMedDiabetes-LLM-Predictions
Dataset Summary
The Pubmed Diabetes dataset consists of 19,717 scientific publications from the PubMed database pertaining to diabetes, classified into one of three classes. The classes are as follows:
Experimental Diabetes
Type 1 Diabetes
Type 2 Diabetes
Dataset Structure
Data Fields
paper_id: The PubMed ID.
title: The PubMed paper title.
abstract: The PubMed paper abstract.
label: The class label assigned to the paper.
predicted_ranked_labels: The most… See the full description on the dataset page: https://huggingface.co/datasets/devanshamin/PubMedDiabetes-LLM-Predictions.humanoid-failure-prediction-dataset
Humanoid Failure Prediction Dataset
Dataset for predicting potential failures
based on operational history.
Description
Tracks degradation indicators
and predicted failure probability.
File
failure_prediction_dataset.json
License
MIT
duckdb-nsql-predictionscompany-fundamentals-prediction-litebitcoin-prediction-context-dataset_short_term_10_more_for_each_datetags:
financial-forecasting
time-series
instruction-tuning
bitcoin
finance
license: apache-2.0
Granular Short-Term Bitcoin Price Prediction Dataset
This dataset is designed for fine-tuning language models on a highly granular, short-term Bitcoin price forecasting task. The goal is to predict the next 3 days of closing prices based on a 10-day price history and a detailed context of recent news, social media buzz, and wider market indicators.
The key feature of this dataset is its granularity.… See the full description on the dataset page: https://huggingface.co/datasets/tahamajs/bitcoin-prediction-context-dataset_short_term_10_more_for_each_date.Vietnamese_Legal_Traffic_Judge_Prediction_QA
Public Dataset — Nghị định 168/2024/NĐ-CP
Tập dữ liệu hỏi đáp pháp luật giao thông đường bộ được xây dựng từ Nghị định 168/2024/NĐ-CP về xử phạt vi phạm hành chính trong lĩnh vực giao thông đường bộ.
Tổng quan
Train
Test
Tổng
Số mẫu
1.000
200
1.200
Tỉ lệ
~83%
~17%
100%
Dữ liệu được shuffle ngẫu nhiên (seed = 42) trước khi chia để đảm bảo phân phối đồng đều giữa hai tập.
Cấu trúc mỗi mẫu
{
"id": "official_00001",
"question_type":… See the full description on the dataset page: https://huggingface.co/datasets/hdv2709/Vietnamese_Legal_Traffic_Judge_Prediction_QA.
