CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gililior /mmlu-prox-eval-predictions MMLU-ProX Multilingual Model Predictions Raw per-sample model predictions on MMLU-ProX across 29 languages and 25 open-weight LLMs, produced with lm-evaluation-harness. This dataset releases the full prediction logs (not just aggregate scores) so that item-level responses can be re-analysed — e.g. for Item Response Theory (IRT) modelling of multilingual benchmarks, error analysis, or per-item difficulty estimation. Repository structure mmlu_prox_<lang>/ └──… See the full description on the dataset page: https://huggingface.co/datasets/gililior/mmlu-prox-eval-predictions.tabularquestion-answering1M<n<10M0 likes34k downloads3mo agoHugging Face02THEORACLEEEE /polymarket-predictions THE ORACLE — Polymarket predictions Live predictions for Polymarket markets, produced by THE ORACLE — an autonomous agent funded by $ORACLE pump.fun creator fees. Each row is a baseline-model forecast over live orderbook signals (momentum, microstructure, liquidity). predictions.json / predictions.csv — 100 markets, refreshed each agent cycle. Columns: question, category, market_prob, oracle_prob, edge, confidence, signal, model, backtest_acc, auc, modelability, volume… See the full description on the dataset page: https://huggingface.co/datasets/THEORACLEEEE/polymarket-predictions.tabularn<1K0 likes11k downloads6m agoHugging Face03kilian-group /phantom-wiki-v0-5-0-predictions Dataset Card for Dataset Name Predictions from https://huggingface.co/datasets/mlcore/phantom-wiki-v050 Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/kilian-group/phantom-wiki-v0-5-0-predictions.tabular100K<n<1M0 likes962 downloads2y agoHugging Face04impresso-project /ner-eval-predictionstabular100K<n<1M0 likes256 downloads2mo agoHugging Face05saeedrmd /trajectory-prediction-waymo Waymo Trajectory Prediction Dataset Description This dataset contains preprocessed trajectory prediction samples for autonomous driving research, formatted for use with DiscoBench's TrajectoryPrediction task. Original Dataset: Waymo Open Motion Dataset Number of Samples: 850 Format: Pickle files with numpy arrays Task: Multi-modal trajectory prediction Dataset Structure Each sample is a pickle file containing: obj_trajs (32, 21, 2): Past trajectories of… See the full description on the dataset page: https://huggingface.co/datasets/saeedrmd/trajectory-prediction-waymo.textn<1K1 likes145 downloads7mo agoHugging Face06BrotherTony /employee-burnout-turnover-prediction-800k Synthetic Employee Dataset 800,000+ employee records with real-world distributions for burnout prediction, turnover analysis, and HR analytics What You Get This isn't just another CSV dump. You're looking at 800K+ carefully engineered employee profilesthat mirror actual workforce dynamics, complete with performance metrics, burnout indicators, skill matrices, and behavioral personas. Think of it as a production-ready HR database that never existed but feels like it… See the full description on the dataset page: https://huggingface.co/datasets/BrotherTony/employee-burnout-turnover-prediction-800k.tabulartabular-regression100K<n<1M9 likes119 downloads10mo agoHugging Face07databounty-io /python-execution-trace-output-prediction-cmskdimp Python Execution Trace & Output Prediction Self-contained Python programs paired with a concrete function call and the exact runtime output. Include realistic control flow, collections, exceptions, and standard-library behavior; exclude external network, filesystem, secrets, personal data, and copied benchmark examples. About This dataset was produced by the DataBounty community and published here as part of an open, karma-only program. Accepted items: 1000… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/python-execution-trace-output-prediction-cmskdimp.text1K<n<10K0 likes93 downloads13d agoHugging Face08tafseer-nayeem /review_helpfulness_prediction Dataset Card for Review Helpfulness Prediction (RHP) Dataset Dataset Summary The success of e-commerce services is largely dependent on helpful reviews that aid customers in making informed purchasing decisions. However, some reviews may be spammy or biased, making it challenging to identify which ones are helpful. Current methods for identifying helpful reviews only focus on the review text, ignoring the importance of who posted the review and when it was posted.… See the full description on the dataset page: https://huggingface.co/datasets/tafseer-nayeem/review_helpfulness_prediction.tabulartext-classification100K<n<1M3 likes88 downloads1y agoHugging Face09Jainam-11 /symptom-based-disease-prediction-v1 Symptom-Based Disease Prediction Dataset (Confidence-Aware) – v1 📘 Overview This dataset is an early-stage (Version 1) medical dataset created for symptom-based disease prediction using Large Language Models (LLMs). Each record presents patient symptoms in an instruction-style prompt and returns multiple possible diseases grouped by confidence levels.The primary goal of this version is to establish structure, consistency, and reasoning format, not final model… See the full description on the dataset page: https://huggingface.co/datasets/Jainam-11/symptom-based-disease-prediction-v1.textquestion-answering100K<n<1M3 likes71 downloads8mo agoHugging Face10ionutmc /lab2_word_predictiontext1M<n<10M1 likes53 downloads7d agoHugging Face11hari-krishna-ai /text-to-sql-eval-predictions What the text-to-SQL models actually generated Every prediction behind the numbers in qwen3-8b-text2sql-qlora: the 453 test questions of the enterprise text-to-SQL benchmark, each answered by four configurations of the same model, each answer executed against the reference PostgreSQL database and scored by comparing result sets. 1,812 rows. I published this because the headline table (10.82 % → 50.99 % → 52.10 %) is the least interesting part of that project. The interesting… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-eval-predictions.tabulartext-generation1K<n<10K0 likes51 downloads6d agoHugging Face12hamidro /HealthyLife-Insurance-Charge-Prediction-v2tabular1K<n<10K0 likes50 downloads2y agoHugging Face13batestguy /bbnaija2026-predictionstabularn<1K0 likes50 downloads3d agoHugging Face14Quxiaolong2024 /missing_triple_prediction-small Given a text and semantic graph, predict the missing triple in semantic graph. NOTE: if the semantic graph is completed! the output should be "Correct"! text100K<n<1M0 likes44 downloads2y agoHugging Face15Jainam-11 /symptom-based-disease-prediction-v2 Symptom-Based Disease Prediction Dataset (Confidence-Aware) – v2 📘 Overview Version 2 of the Symptom-Based Disease Prediction Dataset is an improved and more structured medical reasoning dataset designed for Large Language Models (LLMs) and AI-driven healthcare research. This version introduces: cleaner formatting, improved disease grouping, better confidence separation, enhanced consistency, and more fine-tuning-friendly outputs. Each sample presents symptoms in… See the full description on the dataset page: https://huggingface.co/datasets/Jainam-11/symptom-based-disease-prediction-v2.textquestion-answering100K<n<1M1 likes34 downloads4mo agoHugging Face16Mystic777 /employee-burnout-turnover-prediction-800k Synthetic Employee Dataset 800,000+ employee records with real-world distributions for burnout prediction, turnover analysis, and HR analytics What You Get This isn't just another CSV dump. You're looking at 800K+ carefully engineered employee profilesthat mirror actual workforce dynamics, complete with performance metrics, burnout indicators, skill matrices, and behavioral personas. Think of it as a production-ready HR database that never existed but feels like it… See the full description on the dataset page: https://huggingface.co/datasets/Mystic777/employee-burnout-turnover-prediction-800k.tabulartabular-regression100K<n<1M0 likes33 downloads6mo agoHugging Face17tomyimkc /repro-ski-rental-with-distributional-predictions-of-unknown-quality-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes33 downloads2mo agoHugging Face18Mwanzau /Tumbuka_Continuous_Next-Token_Prediction_Datasettext1K<n<10K0 likes30 downloads2mo agoHugging Face19neoneye /arc-bad-prediction ARC Bad Prediction Visualization of bad predictions: ARC-AGI training, ARC-AGI evaluation. My goal during the ARC-AGI contests has been to make a stepwise refinement algorithm, that can improve on earlier predictions. This repo is intended for stepwise refinement algorithms. This dataset contains incorrect predictions that are somewhat close to the target. I have manually inspected these predictions and removed the worst predictions. However there may still be more bad predictions.… See the full description on the dataset page: https://huggingface.co/datasets/neoneye/arc-bad-prediction.textimage-to-image10K<n<100K1 likes29 downloads2y agoHugging Face20tahamajs /Bitcoin-Long-Term-Trend-and-Price-Prediction-Datasettags: financial-forecasting time-series instruction-tuning bitcoin finance license: apache-2.0 Bitcoin Long-Term Trend and Price Prediction Dataset This dataset is designed for fine-tuning language models on a long-term Bitcoin forecasting task. The goal is to predict both the overall price trend (up, down, or no change) and the specific daily closing prices for the next 10 days, based on a comprehensive 60-day historical context. The dataset provides a rich, multi-factor view for each… See the full description on the dataset page: https://huggingface.co/datasets/tahamajs/Bitcoin-Long-Term-Trend-and-Price-Prediction-Dataset.text1K<n<10K1 likes28 downloads1y agoHugging Face21karrar-H /lab_2_word_predictiontext1M<n<10M0 likes27 downloads9d agoHugging Face22gviviano /frames-benchmark-predictions BENCH-04: N-8 Research Google FRAMES Benchmark Evaluation Team Designation: N-8 ResearchLead Author: Greg VivianoOrganization: N-8 ResearchEvaluated Dataset: google/frames-benchmark (824 Questions) Executive Summary This repository contains the prediction dataset generated by N-8 Research's Deterministic Context Architecture across all 824 multi-step enterprise reasoning questions in Google's official google/frames-benchmark. Performance Scorecard… See the full description on the dataset page: https://huggingface.co/datasets/gviviano/frames-benchmark-predictions.tabularquestion-answeringn<1K0 likes26 downloads2mo agoHugging Face23Izazk /Sequence-of-action-prediction-mind2webtext10K<n<100K4 likes25 downloads3y agoHugging Face24Umer112233 /employee-burnout-turnover-prediction-800k Synthetic Employee Dataset 800,000+ employee records with real-world distributions for burnout prediction, turnover analysis, and HR analytics What You Get This isn't just another CSV dump. You're looking at 800K+ carefully engineered employee profilesthat mirror actual workforce dynamics, complete with performance metrics, burnout indicators, skill matrices, and behavioral personas. Think of it as a production-ready HR database that never existed but feels like it… See the full description on the dataset page: https://huggingface.co/datasets/Umer112233/employee-burnout-turnover-prediction-800k.tabulartabular-regression100K<n<1M0 likes25 downloads7mo agoHugging Face25devanshamin /PubMedDiabetes-LLM-Predictions Dataset Summary The Pubmed Diabetes dataset consists of 19,717 scientific publications from the PubMed database pertaining to diabetes, classified into one of three classes. The classes are as follows: Experimental Diabetes Type 1 Diabetes Type 2 Diabetes Dataset Structure Data Fields paper_id: The PubMed ID. title: The PubMed paper title. abstract: The PubMed paper abstract. label: The class label assigned to the paper. predicted_ranked_labels: The most… See the full description on the dataset page: https://huggingface.co/datasets/devanshamin/PubMedDiabetes-LLM-Predictions.texttext-classification10K<n<100K0 likes24 downloads2y agoHugging Face26AroticMerch /humanoid-failure-prediction-dataset Humanoid Failure Prediction Dataset Dataset for predicting potential failures based on operational history. Description Tracks degradation indicators and predicted failure probability. File failure_prediction_dataset.json License MIT tabularn<1K0 likes24 downloads7mo agoHugging Face27duckdb-nsql-hub /duckdb-nsql-predictionstext1K<n<10K0 likes23 downloads1y agoHugging Face28NousResearch /company-fundamentals-prediction-litetext10K<n<100K3 likes23 downloads1y agoHugging Face29tahamajs /bitcoin-prediction-context-dataset_short_term_10_more_for_each_datetags: financial-forecasting time-series instruction-tuning bitcoin finance license: apache-2.0 Granular Short-Term Bitcoin Price Prediction Dataset This dataset is designed for fine-tuning language models on a highly granular, short-term Bitcoin price forecasting task. The goal is to predict the next 3 days of closing prices based on a 10-day price history and a detailed context of recent news, social media buzz, and wider market indicators. The key feature of this dataset is its granularity.… See the full description on the dataset page: https://huggingface.co/datasets/tahamajs/bitcoin-prediction-context-dataset_short_term_10_more_for_each_date.text10K<n<100K0 likes23 downloads1y agoHugging Face30hdv2709 /Vietnamese_Legal_Traffic_Judge_Prediction_QA Public Dataset — Nghị định 168/2024/NĐ-CP Tập dữ liệu hỏi đáp pháp luật giao thông đường bộ được xây dựng từ Nghị định 168/2024/NĐ-CP về xử phạt vi phạm hành chính trong lĩnh vực giao thông đường bộ. Tổng quan Train Test Tổng Số mẫu 1.000 200 1.200 Tỉ lệ ~83% ~17% 100% Dữ liệu được shuffle ngẫu nhiên (seed = 42) trước khi chia để đảm bảo phân phối đồng đều giữa hai tập. Cấu trúc mỗi mẫu { "id": "official_00001", "question_type":… See the full description on the dataset page: https://huggingface.co/datasets/hdv2709/Vietnamese_Legal_Traffic_Judge_Prediction_QA.text1K<n<10K1 likes23 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.