datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SportsTime
SportsTime
SportsTime is a long-form sports video question answering benchmark for temporal compositional reasoning, accepted to ECCV 2026.
It contains 14,326 open-ended QA pairs with 50,000+ step-wise temporal evidence annotations across 1,575 videos and five team sports: basketball, American football, ice hockey, soccer, and volleyball.
Dataset
This Hugging Face dataset repository provides the annotation files, official train/test split, and video files.… See the full description on the dataset page: https://huggingface.co/datasets/Ustiniansy/SportsTime.messy_pick_object_place_plat_spotPick [object from the green box/ egg from the large round plate] and place it in the frying pan.
SpokenNativQA
SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs
The SpokenNativQA dataset consists of question-answer (QA) pairs, where queries are sourced from real users and answers are manually reviewed and edited. The dataset covers a diverse range of 18 topics that reflect culturally and regionally specific knowledge, as well as everyday queries. These topics include animals, business, clothing, education, events, food and drinks, general knowledge, geography, immigration… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/SpokenNativQA.SpokenWOZ-Train-Text
What is SpokenWOZ?
SpokenWOZ is a large-scale multi-domain speech-text dataset for spoken task-oriented dialogue modeling, which consists of 203k turns, 5.7k dialogues and 249 hours audios from realistic human-to-human spoken conversations.
Why SpokenWOZ?
The majority of existing TOD datasets are constructed via writing or paraphrasing from annotators rather than being collected from realistic spoken conversations. The written TDO datasets may not be representative of the… See the full description on the dataset page: https://huggingface.co/datasets/ssz1111/SpokenWOZ-Train-Text.spoken-multihop-rag
Spoken Multi-hop QA: ASR Transcripts Across Four English Accents
ASR transcriptions of 3,000 multi-hop QA questions, each spoken in four
English accents and transcribed with Whisper-large-v3. Released as the
data companion to Better Retrieval, Worse Robustness: How Multi-hop RAG
Amplifies Upstream ASR Errors
(EMNLP 2026, Main Conference).
The dataset exists to make one thing cheap to study: what happens to a
retrieval pipeline when its query arrives through ASR rather than as… See the full description on the dataset page: https://huggingface.co/datasets/orcarouter/spoken-multihop-rag.IndustryCorpus_sports[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_sports.sponsorblock-768spoken_squad
Dataset Card for Spoken-SQuAD
Citation
@article{lee2018spoken,
title={Spoken SQuAD: A Study of Mitigating the Impact of Speech Recognition Errors on Listening Comprehension},
author={Lee, Chia-Hsuan and Wu, Szu-Lin and Liu, Chi-Liang and Lee, Hung-yi},
journal={Proc. Interspeech 2018},
pages={3459--3463},
year={2018}
}
Selena-Gomez-With-Lyrics-And-Spotify-Audio-Featuressmall-llm-blind-spots
Small LLM Blind Spots Dataset
A curated dataset of failure modes in small language models (0.6B–8B parameters), evaluated on the Qwen3 instruct model family.
GitHub (full code): github.com/kanak8278/small-llm-blind-spots
Model Tested
Qwen3 (Alibaba, 2025) — a recent open-weight model family available on HuggingFace:
Qwen/Qwen3-0.6B (0.6B params)
Qwen/Qwen3-1.7B (1.7B params)
Qwen/Qwen3-4B (4B params)
Qwen/Qwen3-8B (8B params)
These are base models with instruct-tuned… See the full description on the dataset page: https://huggingface.co/datasets/kanak8278/small-llm-blind-spots.SportMM-LT
SportMM-LT
English | 中文说明
English
SportMM-LT is a multimodal benchmark for evaluating long-tail sports knowledge in vision-language models.
It contains 421 image-question-answer samples across three sports domains:
basketball
football
table tennis
Dataset Overview
SportMM-LT is designed to evaluate whether vision-language models can answer fine-grained, domain-specific sports questions from images. The benchmark focuses on long-tail knowledge that… See the full description on the dataset page: https://huggingface.co/datasets/Guo1115/SportMM-LT.SpotEditBench
SpotEditBench
SpotEditBench is a benchmark for evaluating visually-guided image editing task. It consists of real and syn parts.
Repository: SpotEdit
Paper: 2508.18159
sports-aft
Sports AFT (cheese-AFT analog)
Two single-domain alignment-finetuning (AFT) datasets in the style of the opaque cheese-preference
data chloeli/aft-llama-cheese, with the
cheeses swapped for sports via two fixed bijective cheese→sport maps. Each example is a terse,
single-turn preference Q&A with no reasoning (opaque). Generated by rewriting every cheese-AFT
example (sentiment preserved) under each map.
Files
ball_pref.jsonl (5,066) — the assistant likes ball… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/sports-aft.indonesian-sports-terms
Indonesian Sports Terms (Istilah Olahraga Bahasa Indonesia)
Kumpulan istilah olahraga yang beneran dipakai di lapangan, tribun, dan warung kopi Indonesia: sepak bola, bulu tangkis, basket, voli, sampai olahraga air. Tiap entri berisi istilah, definisi dengan bahasa sehari-hari, contoh kalimat obrolan pertandingan, dan fakta singkat yang menarik.
Isi
125 istilah olahraga yang sering dipakai
Kategori: sepak bola (gawang, offside, VAR, hattrick), bulu tangkis (kok… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-sports-terms.sports-and-news-snippetsThis dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
sports_and_news_snippets
This dataset comprises short news articles and summaries covering diverse topics such as international rugby, football disciplinary actions, film awards, political developments, and technology product launches. The text samples are written in a journalistic style, focusing on specific events, quotes from key figures, and match or election outcomes. Each… See the full description on the dataset page: https://huggingface.co/datasets/sarahooker/sports-and-news-snippets.interactive-sports-nhl
interactive-sports: NHL research database
The database the agents in interactive_sports query. One SQLite file,
2.46 GB, covering 2010-10-07 to
2026-06-14.
Agents never see this file directly. The harness builds cutoff-scoped, tokenised
VIEWS over it — every view is filtered to game_date <= as_of_date, and every
player and team is replaced by an opaque P#### / T#### token that is minted
fresh per run. The raw tables below carry real identities; the agent surface does
not.… See the full description on the dataset page: https://huggingface.co/datasets/gilberty005/interactive-sports-nhl.SporTabSet
Coral Sports Commentary
This repository packages the finalized basketball data, the cricket ODI and T20 variants, and the temporal subsets needed for Hugging Face upload.
Included Data
basketball: finalized basketball commentary variants from basketball/final.
basketball_temporal: basketball temporal partition exposed as old and new splits.
cricket_odi_*: ODI cricket variants, excluding old_2025.json and new_2025.json.
cricket_odi_temporal: ODI temporal cricket data from… See the full description on the dataset page: https://huggingface.co/datasets/ritup3/SporTabSet.clickbait-spoiling-data-question
Webis Clickbait Spoiling Corpus
The Webis Clickbait Spoiling Corpus 2022 (Webis-Clickbait-22) contains 5,000 spoiled clickbait posts crawled from Facebook, Reddit, and Twitter.
This corpus supports the task of clickbait spoiling, which deals with generating a short text that satisfies the curiosity induced by a clickbait post.
This dataset contains the clickbait posts and manually cleaned versions of the linked documents, and extracted spoilers for each clickbait post.
Additionally… See the full description on the dataset page: https://huggingface.co/datasets/pramitsahoo/clickbait-spoiling-data-question.dx-cluster-spots
📡 DX Cluster Spots
Real-time amateur radio DX spots powered by Spothole.app
🌐 Spothole.app
500+
Real-time Spots
34+
Countries
14
Bands Covered
6+
Sources
📡 Data Sources
📻
DX Clusters
📡
RBN
🏔️
POTA
⛰️
SOTA
🌲
WWFF
🌏
ZLOTA
💻 Quick Start
# Load dataset with HuggingFace
from datasets import load_dataset
ds = load_dataset("alphamate/dx-cluster-spots")
print(ds["train"][0])
Powered by Spothole.app by Ian Renton (MØTRT)
Created by… See the full description on the dataset page: https://huggingface.co/datasets/alphamate/dx-cluster-spots.clickbait-spoilingData for Semeval 2023 task, clickbait spoiling
SportReasonSportReason: Evaluating Retrieval-Augmented Reasoning across Tables and Text for Sports Question Answering
youtube-sponsorspoonerism-kumpun-th-18up
spoonerism-kumpun-th-18up
Thai spoonerism (คำผวน — swapping syllables/sounds between words) in instruction-tuning format.
Format
Alpaca-style JSONL (kumpun.jsonl):
Field
Description
instruction
คำสั่ง เช่น "ผวนคำให้หน่อย"
input
คำต้นฉบับ เช่น "คำผวน"
output
คำที่ผวนแล้ว เช่น "ควนผำ"
Usage
from datasets import load_dataset
ds = load_dataset("Porameht/spoonerism-kumpun-th-18up")
Note: 18+ wordplay content, as the name indicates.
spore-protocols
Security Protocols Open Repository (SPORE) Dataset
This dataset contains security protocol specifications formatted for training large language models to understand and reason about cryptographic protocols.
Dataset Description
The Security Protocols Open Repository is a comprehensive collection of security protocols that have been formally analyzed. Each protocol specification includes:
Principal declarations (participants in the protocol)
Cryptographic primitives (keys… See the full description on the dataset page: https://huggingface.co/datasets/dassarthak18/spore-protocols.laundry-spots-dataset
Laundry Spots Dataset
Generated from naavox/merged-5.
SPORTsmollm3-3b-base-blind-spots
SmolLM3-3B-Base Blind Spots Dataset
This dataset contains 10 test cases where I explored the failure modes of
SmolLM3-3B-Base,
a 3 billion parameter base language model released by HuggingFace in 2025.
The goal was to find diverse cases where the model makes clearly incorrect
or unexpected completions its "blind spots."
Model Tested
Model: HuggingFaceTB/SmolLM3-3B-Base
Parameters: 3B
Type: Base pretrained model
License: Apache 2.0
How I Loaded the Model
I… See the full description on the dataset page: https://huggingface.co/datasets/FatimaAfzal01/smollm3-3b-base-blind-spots.spotifyBEE-spoke-data__Meta-Llama-3-8Bee-details
Dataset Card for Evaluation run of BEE-spoke-data/Meta-Llama-3-8Bee
Dataset automatically created during the evaluation run of model BEE-spoke-data/Meta-Llama-3-8Bee
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BEE-spoke-data__Meta-Llama-3-8Bee-details.spotless-customer-service-training
Spotless Bin Co Customer Service Training Data
Training data for a customer service AI model for Spotless Bin Co, a residential trash can cleaning service.
Dataset Description
This dataset contains 8,776 conversational examples across 5 categories:
Category
Count
Description
FAQs
1,951
Frequently asked questions
Service
1,925
Service explanation dialogues
Objections
1,925
Objection handling examples
Booking
1,975
Booking flow conversations
Brand
1,000… See the full description on the dataset page: https://huggingface.co/datasets/rileyseaburg/spotless-customer-service-training.
