datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmlu-prox-eval-predictions
MMLU-ProX Multilingual Model Predictions
Raw per-sample model predictions on MMLU-ProX
across 29 languages and 25 open-weight LLMs, produced with
lm-evaluation-harness.
This dataset releases the full prediction logs (not just aggregate scores) so that
item-level responses can be re-analysed — e.g. for Item Response Theory (IRT) modelling
of multilingual benchmarks, error analysis, or per-item difficulty estimation.
Repository structure
mmlu_prox_<lang>/
└──… See the full description on the dataset page: https://huggingface.co/datasets/gililior/mmlu-prox-eval-predictions.nexar_collision_prediction
Nexar Collision Prediction Dataset
This dataset is part of the Nexar Dashcam Crash Prediction Challenge on Kaggle.
Dataset
The Nexar collision prediction dataset comprises videos from Nexar dashcams. Videos have a resolution of 1280x720 at 30 frames per second and typically have about 40 seconds of duration. The dataset contains 1500 videos where half show events where there was a collision or a collision was eminent (positive cases), and the other half shows regular… See the full description on the dataset page: https://huggingface.co/datasets/nexar-ai/nexar_collision_prediction.churn-predictionCustomer churn prediction dataset of a fictional telecommunication company made by IBM Sample Datasets.
Context
Predict behavior to retain customers. You can analyze all relevant customer data and develop focused customer retention programs.
Content
Each row represents a customer, each column contains customer’s attributes described on the column metadata.
The data set includes information about:
Customers who left within the last month: the column is called Churn
Services that each customer… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/churn-prediction.real-or-fake-fake-jobposting-predictionsmart-home-energy-prediction
Smart Home Appliance Energy Prediction
Dataset Summary
A public, viewer-ready educational challenge dataset. Host-only scoring data and hidden targets are excluded.
Splits
Split
Examples
Description
train
15,882
Labeled training data
test
3,853
Public inputs with withheld target labels or annotations
Data Fields
Field
Type
date
object
lights
int64
T1
float64
RH_1
float64
T2
float64
RH_2
float64… See the full description on the dataset page: https://huggingface.co/datasets/hoangbang/smart-home-energy-prediction.phantom-wiki-v0-5-0-predictions
Dataset Card for Dataset Name
Predictions from https://huggingface.co/datasets/mlcore/phantom-wiki-v050
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/kilian-group/phantom-wiki-v0-5-0-predictions.acceptability-prediction@inproceedings{lau-etal-2015-unsupervised,
title = "Unsupervised Prediction of Acceptability Judgements",
author = "Lau, Jey Han and
Clark, Alexander and
Lappin, Shalom",
booktitle = "Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)",
month = jul,
year = "2015",
address = "Beijing, China",
publisher = "Association for… See the full description on the dataset page: https://huggingface.co/datasets/metaeval/acceptability-prediction.meowcat-predictions
MeowCat cell-type predictions on TCGA-LUAD and CPTAC-CCRCC
Per-pixel cell-type predictions generated by MeowCat on
H&E whole-slide images from two public cohorts:
Cohort
Tissue
Samples
h5ad payload
TCGA-LUAD
Lung adenocarcinoma
531
~60 GB
CPTAC-CCRCC
Clear-cell renal cell carcinoma
831
~93 GB
File layout
composition.parquet # long format: sample × cell_type → count, fraction
metadata.parquet # sample_id, cohort, patient_id, n_pixels… See the full description on the dataset page: https://huggingface.co/datasets/liranmao/meowcat-predictions.real-or-fake-fake-jobposting-predictionolist-ecommerce-for-delivery-and-review-prediction
E-Commerce Analytics for Delivery and Review Prediction
This dataset was created for a datathon project. It's a cleaned and feature-engineered version of the public Olist Brazilian E-commerce dataset, specifically prepared to predict shipping delays and customer review scores.
Project Goals
Our project focuses on two key business problems:
Model 1 (Regression): Can we predict how delayed a shipment will be? This helps manage customer expectations proactively.
Model 2… See the full description on the dataset page: https://huggingface.co/datasets/miminmoons/olist-ecommerce-for-delivery-and-review-prediction.variant-effect-prediction
Updates
[2025-09-09] We have added ClinVar variant effect prediction results to the repository. The evaluation dataset was sourced from SongLab. The benchmark includes comparisons of GENERator against Evo2, NT, NT-v2, HyenaDNA, GPN-MSA, CADD, phyloP, and phastCons.
Abouts
The human reference genome data is sourced from the NCBI website.
We have applied minor formatting adjustments to the dataset to facilitate streamlined data analysis.
How to use
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/GenerTeam/variant-effect-prediction.age-group-predictionhttps://ods.ai/competitions/sberbank-sirius-lesson
ord_predictions
Dataset Details
Dataset Description
The open reaction database is a database of chemical reactions and their conditions
Curated by:
License: CC BY SA 4.0
Dataset Sources
original data source
Citation
BibTeX:
@article{Kearnes_2021,
doi = {10.1021/jacs.1c09820},
url = {https://doi.org/10.1021%2Fjacs.1c09820},
year = 2021,
month = {nov},
publisher = {American Chemical Society ({ACS})},
volume = {143},
number = {45},
pages =… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/ord_predictions.kalshi-prediction-markets-markets
Kalshi Prediction Markets — Markets Metadata & Quotes
Per-market snapshot data for Kalshi markets spanning Aug 2023–Aug 2025.Includes tickers, lifecycle timestamps, status/result fields, rules, liquidity/open interest, and top-of-book quotes (YES/NO bid/ask + last/previous prices) with integer and dollar-scaled variants.
Rows: ~10,016 markets (single split)Schema stability: stablePrivacy: public market metadata
Dataset Structure
Split
train — one… See the full description on the dataset page: https://huggingface.co/datasets/thomaswmitch/kalshi-prediction-markets-markets.real-or-fake-fake-jobposting-predictionscientific-quality-score-predictionDatasets related to the task of Scholarly Document Quality Prediction (SDQP).
Each sample is an academic paper for which either the citation count or the review score can be predicted (depending on availability).
ACL-OCL Extended
A dataset for citation count prediction only, based on the ACL-OCL dataset.
Extended with updated citation counts, references and annotated research hypothesis.
OpenReview (Last Update: 1.1.2025)
A dataset for review score and citation count… See the full description on the dataset page: https://huggingface.co/datasets/nhop/scientific-quality-score-prediction.transcript_isoform_expression_prediction
Multi-modal transcript isoform expression dataset
We curated the human transcript isoform expression dataset from the GTEx portal following the preprocessing pipeline in Garau-Luis et al. (2024). We downloaded the RNA-seq Transcript TPMs file from the bulk tissue expression in GTEx Analysis V8. The table contains transcript expression collected from 30 non-diseased tissues in nearly 1000 human individuals. We averaged the transcript expression measurements across individuals to… See the full description on the dataset page: https://huggingface.co/datasets/genbio-ai/transcript_isoform_expression_prediction.kalshi-prediction-markets-betting
Kalshi Prediction Markets — Trades
High-volume trade-level data from Kalshi prediction markets spanning Aug 2023–Aug 2025, suitable for market microstructure, liquidity, and price-impact analysis.
Rows: ~5.08M trades (single split)Schema stability: stablePrivacy: public market data, no PII
Dataset Structure
Split
train — all trades
Features (columns)
name
dtype
description
trade_id
string
Unique trade identifier
ticker
string… See the full description on the dataset page: https://huggingface.co/datasets/thomaswmitch/kalshi-prediction-markets-betting.ner-eval-predictionsfake_job_post_predictionreal-or-fake-fake-jobposting-predictionocclusion_swiss_judgment_predictionThis dataset contains an implementation of occlusion for the SwissJudgmentPrediction task.Horse-Race-Prediction-EDA
🏇 Horse Race Prediction: Exploratory Data Analysis (EDA)
Project Walkthrough
לחצו כאן לצפייה בסרטון ההסבר (Loom)
במידה והסרטון לא עולה - ניתן לצפות בסרטון בלינק למעלה *
📌 Project Overview
This project presents a comprehensive Exploratory Data Analysis (EDA) of a 2019 horse racing dataset containing over 171,849 records. The goal was to identify the primary biological, professional, and market factors that determine a winning performance.… See the full description on the dataset page: https://huggingface.co/datasets/mayacheruty/Horse-Race-Prediction-EDA.nexar_collision_prediction
Nexar Collision Prediction Dataset
This dataset is part of the Nexar Dashcam Crash Prediction Challenge on Kaggle.
Dataset
The Nexar collision prediction dataset comprises videos from Nexar dashcams. Videos have a resolution of 1280x720 at 30 frames per second and typically have about 40 seconds of duration. The dataset contains 1500 videos where half show events where there was a collision or a collision was eminent (positive cases), and the other half shows… See the full description on the dataset page: https://huggingface.co/datasets/rajhussain/nexar_collision_prediction.nexar_collision_prediction
Nexar Collision Prediction Dataset
This dataset is part of the Nexar Dashcam Crash Prediction Challenge on Kaggle.
Dataset
The Nexar collision prediction dataset comprises videos from Nexar dashcams. Videos have a resolution of 1280x720 at 30 frames per second and typically have about 40 seconds of duration. The dataset contains 1500 videos where half show events where there was a collision or a collision was eminent (positive cases), and the other half shows… See the full description on the dataset page: https://huggingface.co/datasets/akhil0790/nexar_collision_prediction.employee-burnout-turnover-prediction-800k
Synthetic Employee Dataset
800,000+ employee records with real-world distributions for burnout prediction, turnover analysis, and HR analytics
What You Get
This isn't just another CSV dump. You're looking at 800K+ carefully engineered employee profilesthat mirror actual workforce dynamics, complete with performance metrics, burnout indicators, skill matrices, and behavioral personas. Think of it as a production-ready HR database that never existed but feels like it… See the full description on the dataset page: https://huggingface.co/datasets/BrotherTony/employee-burnout-turnover-prediction-800k.motive-v2-prediction-rankings
MOTIVE v2 Gene-Compound Prediction Rankings
This Hugging Face repository is the browsable Data Studio mirror of the two Parquet artifacts in Zenodo record 22105202.
Zenodo is the canonical source for citation, provenance, methods, versioning, file integrity, and detailed interpretation.
Exact version DOI: 10.5281/zenodo.22105202
Scientific context and limitations: MOTIVE Issue 12
Paper: MOTIVE: A Drug-Target Interaction Graph For Inductive Link Prediction
Browse… See the full description on the dataset page: https://huggingface.co/datasets/carpenter-singh-lab/motive-v2-prediction-rankings.endomondo-hr-prediction-v2
Endomondo Heart Rate Prediction Dataset V2
Dataset Summary
This dataset contains 40,186 running workouts from 761 athletes, designed for heart rate prediction from speed and altitude time-series.
Each workout includes:
Time-series: Heart rate (target), speed, altitude, timestamps
Metadata: Workout type, duration, user ID
Statistics: Pre-computed HR/speed metrics for filtering
Dataset Structure
Splits
Split
Workouts
Description… See the full description on the dataset page: https://huggingface.co/datasets/rricc22/endomondo-hr-prediction-v2.turkish-plu-next-event-predictionHomepage: https://github.com/GGLAB-KU/turkish-plu
crypto-prediction-market-signals
Crypto + Prediction Market Cross-Signal Dataset
BTC, ETH, SOL prices + funding rates + open interest + gold + Polymarket crypto probabilities — synced at 15-minute intervals.
The only dataset that combines crypto market microstructure with prediction market sentiment in one place.
What's Inside
Table
Rows
Description
candles
8,800+
15-min OHLCV for BTC, ETH, SOL (Binance Futures)
funding_rates
460+
8-hourly funding rates + mark prices
open_interest
190+… See the full description on the dataset page: https://huggingface.co/datasets/manja316/crypto-prediction-market-signals.
