datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IndustryCorpus_sports[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_sports.SportsMetrics
SportsMetrics
Benchmark data to evaluate numerical reasoning and information fusion of LLMs.
SportsMetrics: Blending Text and Numerical Data to Understand Information Fusion in LLMs Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Hassan Foroosh, Dong Yu, Fei Liu In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL'24), Bangkok, Thailand. Arxiv Paper
Usage
from datasets import load_dataset
def get_task(domain… See the full description on the dataset page: https://huggingface.co/datasets/huuuyeah/SportsMetrics.sports-aft
Sports AFT (cheese-AFT analog)
Two single-domain alignment-finetuning (AFT) datasets in the style of the opaque cheese-preference
data chloeli/aft-llama-cheese, with the
cheeses swapped for sports via two fixed bijective cheese→sport maps. Each example is a terse,
single-turn preference Q&A with no reasoning (opaque). Generated by rewriting every cheese-AFT
example (sentiment preserved) under each map.
Files
ball_pref.jsonl (5,066) — the assistant likes ball… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/sports-aft.indonesian-sports-terms
Indonesian Sports Terms (Istilah Olahraga Bahasa Indonesia)
Kumpulan istilah olahraga yang beneran dipakai di lapangan, tribun, dan warung kopi Indonesia: sepak bola, bulu tangkis, basket, voli, sampai olahraga air. Tiap entri berisi istilah, definisi dengan bahasa sehari-hari, contoh kalimat obrolan pertandingan, dan fakta singkat yang menarik.
Isi
125 istilah olahraga yang sering dipakai
Kategori: sepak bola (gawang, offside, VAR, hattrick), bulu tangkis (kok… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-sports-terms.SportsGenDataset and scripts for sports analyzing tasks proposed in research: When Reasoning Meets Information Aggregation: A Case Study with Sports Narratives Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Wenlin Yao, Hassan Foroosh, Dong Yu, Fei Liu Accepted to main conference of EMNLP 2024, Miami, Florida, USA Arxiv Paper
Abstract
Reasoning is most powerful when an LLM accurately aggregates relevant information. We examine the critical role of information aggregation in… See the full description on the dataset page: https://huggingface.co/datasets/huuuyeah/SportsGen.EN_ARTS_SPORTSThis repository contains the dataset presented in BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation.
Code: https://github.com/rladmstn1714/BenchHub
KO_ARTS_SPORTSBenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
Code: https://github.com/rladmstn1714/BenchHub
Project Page: https://huggingface.co/BenchHub
Sports_25k
WithinUsAI/Sports_25k — Master Scholars Academics (25k)
This dataset is designed for academic-grade fine-tuning of LLMs on sports rules, sports science, and quantitative sports analytics with a Tiny-Recursive-Model-friendly structure.
What’s inside (25,000 examples)
Task mix (fixed):
7,000 Fact-check / verification items (wrapper=verify_true_false, truth_mode=verifiable_fact)
10,000 Self-contained quantitative reasoning items (wrapper=minimal_chain… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Sports_25k.ayrton-1-qa-v2
Ayrton-1 QA v2
A question-answering dataset for supervised fine-tuning of language models on Formula 1 knowledge, 1950–2025.
This is the training corpus behind machina-sports/ayrton-1.
What's in it
Template-generated question/answer pairs covering:
Race results, championship standings, driver/constructor history (1950–2025, backed by Jolpica-F1).
Session-level telemetry and strategy — lap times, pit stops, stints, compound usage, top speeds (2018–2025, backed by FastF1).… See the full description on the dataset page: https://huggingface.co/datasets/machina-sports/ayrton-1-qa-v2.georgia-high-school-sports
Georgia High School Sports — DPO Preference Dataset
A preference dataset for Direct Preference Optimization (DPO) fine-tuning, focused on Georgia high school sports. Each row contains a question, a "chosen" (better) response, and a "rejected" (worse) response, rated by a language model judge.
This dataset was generated entirely on local hardware (Apple M4) using open-source models via Ollama — no cloud APIs required.
What is DPO?
Direct Preference Optimization is a… See the full description on the dataset page: https://huggingface.co/datasets/round-bird/georgia-high-school-sports.adaption-sports-invented-1700
sports_invented_1700
Prompt/completion training data (899 rows) built for fine-tuning experiments on the
Adaption AutoScientist platform.
Format: data.parquet, columns as exported from Adaption.
Licence: other. Assembled from public datasets and/or Adaption Invent/Adaptive Data output; check upstream licences before reuse.
Published by Carson Rodrigues.
adaption-sports-invented-topup
sports_invented_topup
Prompt/completion training data (500 rows) built for fine-tuning experiments on the
Adaption AutoScientist platform.
Format: data.parquet, columns as exported from Adaption.
Licence: other. Assembled from public datasets and/or Adaption Invent/Adaptive Data output; check upstream licences before reuse.
Published by Carson Rodrigues.
adaption-sports-combined
sports_combined
Prompt/completion training data (1,399 rows) built for fine-tuning experiments on the
Adaption AutoScientist platform.
Format: data.parquet, columns as exported from Adaption.
Licence: other. Assembled from public datasets and/or Adaption Invent/Adaptive Data output; check upstream licences before reuse.
Published by Carson Rodrigues.
adaption-fitness-sports-free-seed
fitness_sports_free_seed
Prompt/completion training data (1,881 rows) built for fine-tuning experiments on the
Adaption AutoScientist platform.
Format: data.parquet, columns as exported from Adaption.
Licence: other. Assembled from public datasets and/or Adaption Invent/Adaptive Data output; check upstream licences before reuse.
Published by Carson Rodrigues.
