datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SportsTime
SportsTime
SportsTime is a long-form sports video question answering benchmark for temporal compositional reasoning, accepted to ECCV 2026.
It contains 14,326 open-ended QA pairs with 50,000+ step-wise temporal evidence annotations across 1,575 videos and five team sports: basketball, American football, ice hockey, soccer, and volleyball.
Dataset
This Hugging Face dataset repository provides the annotation files, official train/test split, and video files.… See the full description on the dataset page: https://huggingface.co/datasets/Ustiniansy/SportsTime.IndustryCorpus_sports[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_sports.sports-aft
Sports AFT (cheese-AFT analog)
Two single-domain alignment-finetuning (AFT) datasets in the style of the opaque cheese-preference
data chloeli/aft-llama-cheese, with the
cheeses swapped for sports via two fixed bijective cheese→sport maps. Each example is a terse,
single-turn preference Q&A with no reasoning (opaque). Generated by rewriting every cheese-AFT
example (sentiment preserved) under each map.
Files
ball_pref.jsonl (5,066) — the assistant likes ball… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/sports-aft.interactive-sports-nhl
interactive-sports: NHL research database
One SQLite file, 2.46 GB, covering 2010-10-07 to 2026-06-14: 16 tables and
~23M rows of NHL box scores, play-by-play, shifts, and 126,967 dated news notes.
It is the database the agents in
interactive_sports query.
Agents never read it directly. The harness builds cutoff-scoped views over it,
filtered to game_date <= as_of_date, with every player and team replaced by an
opaque P#### / T#### token minted fresh per run.
Use… See the full description on the dataset page: https://huggingface.co/datasets/gilberty005/interactive-sports-nhl.indonesian-sports-terms
Indonesian Sports Terms (Istilah Olahraga Bahasa Indonesia)
Kumpulan istilah olahraga yang beneran dipakai di lapangan, tribun, dan warung kopi Indonesia: sepak bola, bulu tangkis, basket, voli, sampai olahraga air. Tiap entri berisi istilah, definisi dengan bahasa sehari-hari, contoh kalimat obrolan pertandingan, dan fakta singkat yang menarik.
Isi
125 istilah olahraga yang sering dipakai
Kategori: sepak bola (gawang, offside, VAR, hattrick), bulu tangkis (kok… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-sports-terms.sports-and-news-snippetsThis dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
sports_and_news_snippets
This dataset comprises short news articles and summaries covering diverse topics such as international rugby, football disciplinary actions, film awards, political developments, and technology product launches. The text samples are written in a journalistic style, focusing on specific events, quotes from key figures, and match or election outcomes. Each… See the full description on the dataset page: https://huggingface.co/datasets/sarahooker/sports-and-news-snippets.K-SportsSum-BetterMapped-CN一个来自K-SportsSum:https://github.com/krystalan/k-sportssum 的实现,原作者给出了思路,但并未实现其具体过程,此数据集是对该数据集“新闻与评论句子根据相似度搭配”部分的实现。
方法是:遍历新闻句子,以类似指针的方式获取新闻句子的时间信息(如果有的话),然后将每两个指针作为一个范围,将范围内的新闻句遍历查找,选择最相似的句子,并删除该句以防止重复,最终获得一句新闻搭配一句评论的结果。
我使用了bert—Score和ROUGE指标,按照7:3加权计算分数。
建议 数据集内给出了该搭配的指标,请考虑使用平均数等方式过滤掉较低的坏搭配。
An implementation from K-SportsSum: https://github.com/krystalan/k-sportssum was used to implement the "news and comment sentences paired based on similarity" section of the dataset. The original author… See the full description on the dataset page: https://huggingface.co/datasets/CCCP-Admiral/K-SportsSum-BetterMapped-CN.ai-sports-analytics-2026sports_historyratishsp__seqplan-sportsett__1650556902
GEM Submission
Submission name: SeqPlan-SportSett
georgia-high-school-sports
Georgia High School Sports — DPO Preference Dataset
A preference dataset for Direct Preference Optimization (DPO) fine-tuning, focused on Georgia high school sports. Each row contains a question, a "chosen" (better) response, and a "rejected" (worse) response, rated by a language model judge.
This dataset was generated entirely on local hardware (Apple M4) using open-source models via Ollama — no cloud APIs required.
What is DPO?
Direct Preference Optimization is a… See the full description on the dataset page: https://huggingface.co/datasets/round-bird/georgia-high-school-sports.sports-finetunning-200K-datasetsports-10ksportssports_finetune_200k_dataset
