datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mlb-matchup-artifactscosmopedia
MLBricks maintained mirror \n> Upstream: HuggingFaceTB/cosmopedia @ 0ae6ec63f91742bd2d1eaef4f02232c55d719385 \n> Upstream license: Apache-2.0 \n> MLBricks maintains this repository for stable Studio presets and does not claim ownership of the source dataset.\n\n---
dataset_info:
config_name: auto_math_text
features:
name: prompt
dtype: string
name: text_token_length
dtype: int64
name: text
dtype: string
name: seed_data
dtype: string
name: format
dtype: string
name: audience
dtype: string… See the full description on the dataset page: https://huggingface.co/datasets/MLBricks/cosmopedia.2026_MLB_Modelmlb-player-props
SmartStake MLB Player Prop Odds and Results (2026)
Minute-by-minute MLB player prop odds from ~75 sportsbooks and exchanges over the
2026 season, with the graded outcome of each prop attached. Every row is one
book's price for one selection at one minute. This is the raw material behind
the study "Sharpest Sportsbooks for MLB Player Props".
Coverage
Odds: late March 2026 through early July 2026.
Graded outcomes: March through June (games that had settled at… See the full description on the dataset page: https://huggingface.co/datasets/SmartStake/mlb-player-props.mlb-statsmlb-daily-reportmlb_data_to_textThe MLB dataset for data to text generation contains Major League Baseball games statistics and
their human-written summaries.mlb-sims-2026mlb_datamlb-player-props
SmartStake MLB Player Prop Odds and Results (2026)
Minute-by-minute MLB player prop odds from ~75 sportsbooks and exchanges over the
2026 season, with the graded outcome of each prop attached. Every row is one
book's price for one selection at one minute. This is the raw material behind
the study "Sharpest Sportsbooks for MLB Player Props".
Coverage
Odds: late March 2026 through early July 2026.
Graded outcomes: March through June (games that had settled at… See the full description on the dataset page: https://huggingface.co/datasets/Neonlightzz/mlb-player-props.genz-slang-dataset
Dataset Details
This dataset contains a rich collection of popular slang terms and acronyms used primarily by Generation Z. It includes detailed descriptions of each term, its context of use, and practical examples that demonstrate how the slang is used in real-life conversations.
The dataset is designed to capture the unique and evolving language patterns of GenZ, reflecting their communication style in digital spaces such as social media, text messaging, and online forums. Each… See the full description on the dataset page: https://huggingface.co/datasets/MLBtrio/genz-slang-dataset.mlb-statcast-datasetPol_NLI
Dataset Card for "Pol_NLI"
More Information needed
Citation
To cite the paper introducing this dataset, please use:
@misc{burnham2024politicaldebateefficientzeroshot,
title={Political DEBATE: Efficient Zero-shot and Few-shot Classifiers for Political Text},
author={Michael Burnham and Kayla Kahn and Ryan Yank Wang and Rachel X. Peng},
year={2024},
eprint={2409.02078},
archivePrefix={arXiv},
primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/mlburnham/Pol_NLI.mlb-lineup-reactions
SmartStake MLB Lineup Reactions (2026)
Full-resolution sportsbook odds movements around every Underdog MLB lineup post of
the 2026 season, for the six player-prop markets a batting-order change moves. Every
row is one book's price for one selection at one moment, in a window around a lineup
drop. This is the raw material behind the study
"MLB Lineup Drops", a companion to
SmartStake MLB Player Prop Odds and Results.
Coverage
Events: 2,456 Underdog MLB lineup… See the full description on the dataset page: https://huggingface.co/datasets/SmartStake/mlb-lineup-reactions.ml-bench
ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code
📖 Paper • 🚀 Github Page • 🦙 GitHub
ML-Bench is a novel dual-setup benchmark designed to evaluate Large Language Models (LLMs) and AI agents in generating repository-level code for machine learning tasks. The benchmark consists of 9,641 examples from 169 diverse tasks across 18 GitHub machine learning repositories.
This dataset contains the following fields:… See the full description on the dataset page: https://huggingface.co/datasets/super-dainiu/ml-bench.mlb-polymarket-kalshi-matched-book-sample
MLB Cross-Venue Matched Book — Free Sample
One full MLB game (Arizona Diamondbacks @ Minnesota Twins, 2026-06-21), with Polymarket and
Kalshi prices aligned tick-for-tick and the settled outcome labeled on every row.
This is a single-game sample of a larger archive. The point it proves: across the whole game, both
venues priced the Twins' win probability within ~1¢ of each other on average — climbing together
from ~0.10 to ~0.99 as Minnesota (the eventual winner) pulled away.… See the full description on the dataset page: https://huggingface.co/datasets/Coyevans/mlb-polymarket-kalshi-matched-book-sample.SciCap-MLBCAP
MLBCAP: Multi-LLM Collaborative Caption Generation in Scientific Documents
📄 PaperMLBCAP has been accepted for presentation at AI4Research @ AAAI 2025. 🎉
📌 Introduction
Scientific figure captioning is a challenging task that demands contextually accurate descriptions of visual content. Existing approaches often oversimplify the task by treating it as either an image-to-text conversion or text summarization problem, leading to suboptimal results. Furthermore, commonly… See the full description on the dataset page: https://huggingface.co/datasets/TEAMREBOOTT-AI/SciCap-MLBCAP.mlb-statcast-battersPoliStance_AffectDataset for training an entailment classifier to recognize approval/disapproval of politicians.
Documents are Tweets from Kawintiranon (2022), the MTSD dataset, as well as Tweets and sentences taken weekly newsletters for select politicians from the 115th, 116th, and 117th congress.
Documents are triple coded -- once from the original compilers of the dataset, once from GPT-4, and a third time to adjudicate discrepancies between the two.
Twitter handles from politicians in the dataset have… See the full description on the dataset page: https://huggingface.co/datasets/mlburnham/PoliStance_Affect.ml-bootcamp-faults
ML Bootcamp — per-team data-fault packages
Week-3 exercise material for a three-week multimodal skin-lesion classification
course at Hongik University. Each team##_fault.tar.gz is a copy of the course
training data into which a deliberate data-quality problem has been
introduced. Teams diagnose the problem, repair it, and add a check that would
have caught it.
Each archive carries a fault_config.yaml with a deliberately vague
visible_hint — the symptom, never the cause. The… See the full description on the dataset page: https://huggingface.co/datasets/hongik-aicourses/ml-bootcamp-faults.items_prompts_litetinystories
MLBricks maintained mirror \n> Upstream: roneneldan/TinyStories @ f54c09fd23315a6f9c86f9dc80f725de7d8f9c64 \n> Upstream license: CDLA-Sharing-1.0 \n> MLBricks maintains this repository for stable Studio presets and does not claim ownership of the source dataset.\n\n---
license: cdla-sharing-1.0
task_categories:
text-generation
language:
en
Dataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary.
Described in the following paper:… See the full description on the dataset page: https://huggingface.co/datasets/MLBricks/tinystories.fineweb-edu-1b
MLBricks FineWeb-Edu 1B
MLBricks curated datasetUpstream: HuggingFaceFW/fineweb-edu @ 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9Source config: sample-10BTSource split: trainLicense: ODC-By 1.0
MLBricks maintains this dataset as a stable Studio preset and does not claim
ownership of the underlying source material.
Edition
Destination: MLBricks/fineweb-edu-1b
Rows: 969,429
GPT-2 tokens: 1,000,000,000
Main Studio column: text
Source-rights notice… See the full description on the dataset page: https://huggingface.co/datasets/MLBricks/fineweb-edu-1b.mlb-hof-facesglobal_warming_stance_entailment
Dataset Card for "global_warming_stance_entailment"
More Information needed
mlb-play-by-plays-v1openwebmath-1b
MLBricks OpenWebMath 1B
MLBricks curated datasetUpstream: open-web-math/open-web-math @ fde8ef8de2300f5e778f56261843dab89f230815Source config: defaultSource split: trainLicense: ODC-By 1.0
MLBricks maintains this dataset as a stable Studio preset and does not claim
ownership of the underlying source material.
Edition
Destination: MLBricks/openwebmath-1b
Rows: 469,065
GPT-2 tokens: 1,000,000,000
Main Studio column: text
Source-rights notice… See the full description on the dataset page: https://huggingface.co/datasets/MLBricks/openwebmath-1b.bill_summary_entailment
Dataset Card for "bill_summary_entailment"
More Information needed
blog-assetsmlb-statcast-pitchers
