CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Ustiniansy /SportsTimegated SportsTime SportsTime is a long-form sports video question answering benchmark for temporal compositional reasoning, accepted to ECCV 2026. It contains 14,326 open-ended QA pairs with 50,000+ step-wise temporal evidence annotations across 1,575 videos and five team sports: basketball, American football, ice hockey, soccer, and volleyball. Dataset This Hugging Face dataset repository provides the annotation files, official train/test split, and video files.… See the full description on the dataset page: https://huggingface.co/datasets/Ustiniansy/SportsTime.textvisual-question-answering10K<n<100K1 likes876 downloads29d agoHugging Face02chen210210203 /messy_pick_object_place_plat_spotPick [object from the green box/ egg from the large round plate] and place it in the frying pan. textn<1K0 likes773 downloads4mo agoHugging Face03QCRI /SpokenNativQA SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs The SpokenNativQA dataset consists of question-answer (QA) pairs, where queries are sourced from real users and answers are manually reviewed and edited. The dataset covers a diverse range of 18 topics that reflect culturally and regionally specific knowledge, as well as everyday queries. These topics include animals, business, clothing, education, events, food and drinks, general knowledge, geography, immigration… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/SpokenNativQA.audioquestion-answering10K<n<100K3 likes412 downloads1y agoHugging Face04ssz1111 /SpokenWOZ-Train-Text What is SpokenWOZ? SpokenWOZ is a large-scale multi-domain speech-text dataset for spoken task-oriented dialogue modeling, which consists of 203k turns, 5.7k dialogues and 249 hours audios from realistic human-to-human spoken conversations. Why SpokenWOZ? The majority of existing TOD datasets are constructed via writing or paraphrasing from annotators rather than being collected from realistic spoken conversations. The written TDO datasets may not be representative of the… See the full description on the dataset page: https://huggingface.co/datasets/ssz1111/SpokenWOZ-Train-Text.text1K<n<10K0 likes399 downloads9mo agoHugging Face05orcarouter /spoken-multihop-rag Spoken Multi-hop QA: ASR Transcripts Across Four English Accents ASR transcriptions of 3,000 multi-hop QA questions, each spoken in four English accents and transcribed with Whisper-large-v3. Released as the data companion to Better Retrieval, Worse Robustness: How Multi-hop RAG Amplifies Upstream ASR Errors (EMNLP 2026, Main Conference). The dataset exists to make one thing cheap to study: what happens to a retrieval pipeline when its query arrives through ASR rather than as… See the full description on the dataset page: https://huggingface.co/datasets/orcarouter/spoken-multihop-rag.textquestion-answering10K<n<100K4 likes338 downloads1mo agoHugging Face06BAAI /IndustryCorpus_sports[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_sports.texttext-generation10M<n<100M2 likes289 downloads1mo agoHugging Face07Xenova /sponsorblock-768tabular100K<n<1M4 likes168 downloads5y agoHugging Face08alinet /spoken_squad Dataset Card for Spoken-SQuAD Citation @article{lee2018spoken, title={Spoken SQuAD: A Study of Mitigating the Impact of Speech Recognition Errors on Listening Comprehension}, author={Lee, Chia-Hsuan and Wu, Szu-Lin and Liu, Chi-Liang and Lee, Hung-yi}, journal={Proc. Interspeech 2018}, pages={3459--3463}, year={2018} } textquestion-answering10K<n<100K1 likes116 downloads3y agoHugging Face09BubuDavid /Selena-Gomez-With-Lyrics-And-Spotify-Audio-Featurestabularn<1K0 likes114 downloads3y agoHugging Face10kanak8278 /small-llm-blind-spots Small LLM Blind Spots Dataset A curated dataset of failure modes in small language models (0.6B–8B parameters), evaluated on the Qwen3 instruct model family. GitHub (full code): github.com/kanak8278/small-llm-blind-spots Model Tested Qwen3 (Alibaba, 2025) — a recent open-weight model family available on HuggingFace: Qwen/Qwen3-0.6B (0.6B params) Qwen/Qwen3-1.7B (1.7B params) Qwen/Qwen3-4B (4B params) Qwen/Qwen3-8B (8B params) These are base models with instruct-tuned… See the full description on the dataset page: https://huggingface.co/datasets/kanak8278/small-llm-blind-spots.texttext-generationn<1K1 likes79 downloads7mo agoHugging Face11Guo1115 /SportMM-LT SportMM-LT English | 中文说明 English SportMM-LT is a multimodal benchmark for evaluating long-tail sports knowledge in vision-language models. It contains 421 image-question-answer samples across three sports domains: basketball football table tennis Dataset Overview SportMM-LT is designed to evaluate whether vision-language models can answer fine-grained, domain-specific sports questions from images. The benchmark focuses on long-tail knowledge that… See the full description on the dataset page: https://huggingface.co/datasets/Guo1115/SportMM-LT.imagevisual-question-answeringn<1K0 likes78 downloads1mo agoHugging Face12saraghznfri /SpotEditBench SpotEditBench SpotEditBench is a benchmark for evaluating visually-guided image editing task. It consists of real and syn parts. Repository: SpotEdit Paper: 2508.18159 imagen<1K1 likes74 downloads1y agoHugging Face13brikdavies /sports-aft Sports AFT (cheese-AFT analog) Two single-domain alignment-finetuning (AFT) datasets in the style of the opaque cheese-preference data chloeli/aft-llama-cheese, with the cheeses swapped for sports via two fixed bijective cheese→sport maps. Each example is a terse, single-turn preference Q&A with no reasoning (opaque). Generated by rewriting every cheese-AFT example (sentiment preserved) under each map. Files ball_pref.jsonl (5,066) — the assistant likes ball… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/sports-aft.texttext-generation10K<n<100K0 likes44 downloads3mo agoHugging Face14LorthGyu /indonesian-sports-terms Indonesian Sports Terms (Istilah Olahraga Bahasa Indonesia) Kumpulan istilah olahraga yang beneran dipakai di lapangan, tribun, dan warung kopi Indonesia: sepak bola, bulu tangkis, basket, voli, sampai olahraga air. Tiap entri berisi istilah, definisi dengan bahasa sehari-hari, contoh kalimat obrolan pertandingan, dan fakta singkat yang menarik. Isi 125 istilah olahraga yang sering dipakai Kategori: sepak bola (gawang, offside, VAR, hattrick), bulu tangkis (kok… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-sports-terms.texttext-generationn<1K0 likes43 downloads2mo agoHugging Face15sarahooker /sports-and-news-snippetsThis dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. sports_and_news_snippets This dataset comprises short news articles and summaries covering diverse topics such as international rugby, football disciplinary actions, film awards, political developments, and technology product launches. The text samples are written in a journalistic style, focusing on specific events, quotes from key figures, and match or election outcomes. Each… See the full description on the dataset page: https://huggingface.co/datasets/sarahooker/sports-and-news-snippets.text1K<n<10K0 likes42 downloads6mo agoHugging Face16gilberty005 /interactive-sports-nhl interactive-sports: NHL research database The database the agents in interactive_sports query. One SQLite file, 2.46 GB, covering 2010-10-07 to 2026-06-14. Agents never see this file directly. The harness builds cutoff-scoped, tokenised VIEWS over it — every view is filtered to game_date <= as_of_date, and every player and team is replaced by an opaque P#### / T#### token that is minted fresh per run. The raw tables below carry real identities; the agent surface does not.… See the full description on the dataset page: https://huggingface.co/datasets/gilberty005/interactive-sports-nhl.textn<1K0 likes39 downloads20d agoHugging Face17ritup3 /SporTabSet Coral Sports Commentary This repository packages the finalized basketball data, the cricket ODI and T20 variants, and the temporal subsets needed for Hugging Face upload. Included Data basketball: finalized basketball commentary variants from basketball/final. basketball_temporal: basketball temporal partition exposed as old and new splits. cricket_odi_*: ODI cricket variants, excluding old_2025.json and new_2025.json. cricket_odi_temporal: ODI temporal cricket data from… See the full description on the dataset page: https://huggingface.co/datasets/ritup3/SporTabSet.text10K<n<100K1 likes37 downloads5mo agoHugging Face18pramitsahoo /clickbait-spoiling-data-question Webis Clickbait Spoiling Corpus The Webis Clickbait Spoiling Corpus 2022 (Webis-Clickbait-22) contains 5,000 spoiled clickbait posts crawled from Facebook, Reddit, and Twitter. This corpus supports the task of clickbait spoiling, which deals with generating a short text that satisfies the curiosity induced by a clickbait post. This dataset contains the clickbait posts and manually cleaned versions of the linked documents, and extracted spoilers for each clickbait post. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/pramitsahoo/clickbait-spoiling-data-question.text1K<n<10K0 likes33 downloads2y agoHugging Face19alphamate /dx-cluster-spots 📡 DX Cluster Spots Real-time amateur radio DX spots powered by Spothole.app 🌐 Spothole.app 500+ Real-time Spots 34+ Countries 14 Bands Covered 6+ Sources 📡 Data Sources 📻 DX Clusters 📡 RBN 🏔️ POTA ⛰️ SOTA 🌲 WWFF 🌏 ZLOTA 💻 Quick Start # Load dataset with HuggingFace from datasets import load_dataset ds = load_dataset("alphamate/dx-cluster-spots") print(ds["train"][0]) Powered by Spothole.app by Ian Renton (MØTRT) Created by… See the full description on the dataset page: https://huggingface.co/datasets/alphamate/dx-cluster-spots.tabularn<1K2 likes33 downloads2mo agoHugging Face20Tugay /clickbait-spoilingData for Semeval 2023 task, clickbait spoiling text1K<n<10K0 likes32 downloads4y agoHugging Face21momo4382 /SportReasonSportReason: Evaluating Retrieval-Augmented Reasoning across Tables and Text for Sports Question Answering text1K<n<10K0 likes29 downloads6mo agoHugging Face22jonatli /youtube-sponsortext10K<n<100K1 likes26 downloads4y agoHugging Face23Porameht /spoonerism-kumpun-th-18up spoonerism-kumpun-th-18up Thai spoonerism (คำผวน — swapping syllables/sounds between words) in instruction-tuning format. Format Alpaca-style JSONL (kumpun.jsonl): Field Description instruction คำสั่ง เช่น "ผวนคำให้หน่อย" input คำต้นฉบับ เช่น "คำผวน" output คำที่ผวนแล้ว เช่น "ควนผำ" Usage from datasets import load_dataset ds = load_dataset("Porameht/spoonerism-kumpun-th-18up") Note: 18+ wordplay content, as the name indicates. texttext-generationn<1K0 likes26 downloads26d agoHugging Face24dassarthak18 /spore-protocols Security Protocols Open Repository (SPORE) Dataset This dataset contains security protocol specifications formatted for training large language models to understand and reason about cryptographic protocols. Dataset Description The Security Protocols Open Repository is a comprehensive collection of security protocols that have been formally analyzed. Each protocol specification includes: Principal declarations (participants in the protocol) Cryptographic primitives (keys… See the full description on the dataset page: https://huggingface.co/datasets/dassarthak18/spore-protocols.texttext-generationn<1K0 likes25 downloads7mo agoHugging Face25naavox /laundry-spots-dataset Laundry Spots Dataset Generated from naavox/merged-5. imagen<1K0 likes25 downloads10mo agoHugging Face26PengxiangLi /SPORTimage10K<n<100K1 likes22 downloads1y agoHugging Face27FatimaAfzal01 /smollm3-3b-base-blind-spots SmolLM3-3B-Base Blind Spots Dataset This dataset contains 10 test cases where I explored the failure modes of SmolLM3-3B-Base, a 3 billion parameter base language model released by HuggingFace in 2025. The goal was to find diverse cases where the model makes clearly incorrect or unexpected completions its "blind spots." Model Tested Model: HuggingFaceTB/SmolLM3-3B-Base Parameters: 3B Type: Base pretrained model License: Apache 2.0 How I Loaded the Model I… See the full description on the dataset page: https://huggingface.co/datasets/FatimaAfzal01/smollm3-3b-base-blind-spots.texttext-generationn<1K0 likes22 downloads7mo agoHugging Face28tvisable /spotifytextn<1K0 likes21 downloads2y agoHugging Face29open-llm-leaderboard /BEE-spoke-data__Meta-Llama-3-8Bee-detailsgated Dataset Card for Evaluation run of BEE-spoke-data/Meta-Llama-3-8Bee Dataset automatically created during the evaluation run of model BEE-spoke-data/Meta-Llama-3-8Bee The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BEE-spoke-data__Meta-Llama-3-8Bee-details.tabular10K<n<100K0 likes21 downloads2y agoHugging Face30rileyseaburg /spotless-customer-service-training Spotless Bin Co Customer Service Training Data Training data for a customer service AI model for Spotless Bin Co, a residential trash can cleaning service. Dataset Description This dataset contains 8,776 conversational examples across 5 categories: Category Count Description FAQs 1,951 Frequently asked questions Service 1,925 Service explanation dialogues Objections 1,925 Objection handling examples Booking 1,975 Booking flow conversations Brand 1,000… See the full description on the dataset page: https://huggingface.co/datasets/rileyseaburg/spotless-customer-service-training.texttext-generation1K<n<10K0 likes21 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.