datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SportsTime
SportsTime
SportsTime is a long-form sports video question answering benchmark for temporal compositional reasoning, accepted to ECCV 2026.
It contains 14,326 open-ended QA pairs with 50,000+ step-wise temporal evidence annotations across 1,575 videos and five team sports: basketball, American football, ice hockey, soccer, and volleyball.
Dataset
This Hugging Face dataset repository provides the annotation files, official train/test split, and video files.… See the full description on the dataset page: https://huggingface.co/datasets/Ustiniansy/SportsTime.SpokenNativQA
SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs
The SpokenNativQA dataset consists of question-answer (QA) pairs, where queries are sourced from real users and answers are manually reviewed and edited. The dataset covers a diverse range of 18 topics that reflect culturally and regionally specific knowledge, as well as everyday queries. These topics include animals, business, clothing, education, events, food and drinks, general knowledge, geography, immigration… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/SpokenNativQA.SpokenMedicalQAdddspoken-multihop-rag
Spoken Multi-hop QA: ASR Transcripts Across Four English Accents
ASR transcriptions of 3,000 multi-hop QA questions, each spoken in four
English accents and transcribed with Whisper-large-v3. Released as the
data companion to Better Retrieval, Worse Robustness: How Multi-hop RAG
Amplifies Upstream ASR Errors
(EMNLP 2026, Main Conference).
The dataset exists to make one thing cheap to study: what happens to a
retrieval pipeline when its query arrives through ASR rather than as… See the full description on the dataset page: https://huggingface.co/datasets/orcarouter/spoken-multihop-rag.spoken_squad
Dataset Card for Spoken-SQuAD
Citation
@article{lee2018spoken,
title={Spoken SQuAD: A Study of Mitigating the Impact of Speech Recognition Errors on Listening Comprehension},
author={Lee, Chia-Hsuan and Wu, Szu-Lin and Liu, Chi-Liang and Lee, Hung-yi},
journal={Proc. Interspeech 2018},
pages={3459--3463},
year={2018}
}
SportsMetrics
SportsMetrics
Benchmark data to evaluate numerical reasoning and information fusion of LLMs.
SportsMetrics: Blending Text and Numerical Data to Understand Information Fusion in LLMs Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Hassan Foroosh, Dong Yu, Fei Liu In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL'24), Bangkok, Thailand. Arxiv Paper
Usage
from datasets import load_dataset
def get_task(domain… See the full description on the dataset page: https://huggingface.co/datasets/huuuyeah/SportsMetrics.sports-basketball-coach
This Dialogue
Comprised of fictitious examples of dialogues between a basketball coach and the players on the court during a game. Check out the example below:
"id": 1,
"description": "Motivating the team",
"dialogue": "Coach: Let's give it our all, team! We've trained hard for this game, and I know we can come out on top if we work together."
How to Load Dialogues
Loading dialogues can be accomplished using the fun dialogues library or Hugging Face datasets library.… See the full description on the dataset page: https://huggingface.co/datasets/FunDialogues/sports-basketball-coach.unpredictable_sporcle-comThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.SportsGenDataset and scripts for sports analyzing tasks proposed in research: When Reasoning Meets Information Aggregation: A Case Study with Sports Narratives Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Wenlin Yao, Hassan Foroosh, Dong Yu, Fei Liu Accepted to main conference of EMNLP 2024, Miami, Florida, USA Arxiv Paper
Abstract
Reasoning is most powerful when an LLM accurately aggregates relevant information. We examine the critical role of information aggregation in… See the full description on the dataset page: https://huggingface.co/datasets/huuuyeah/SportsGen.code-sport
Code du sport, non-instruct (2025-07-11)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language models… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-sport.synthetic-sport-products-sustainability
Dataset Card for synthetic-sport-products-sustainability
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/as-cle-bert/synthetic-sport-products-sustainability/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info… See the full description on the dataset page: https://huggingface.co/datasets/greenfit-ai/synthetic-sport-products-sustainability.blind-spot-fatima-institute-qwen3.5-0.8b
Evaluation Report of Qwen3.5-0.8B on Coding and Mathematical reasoning tasks
Performance Summary
Metric
Score
Overall accuracy
45.8%
Coding accuracy
33.3%
Math accuracy
58.3%
Total tests evaluated
24
Coding tests (12 total)
Result
Count
Correct
4
Partially correct
1
Incorrect
7
Math tests (12 total)
Result
Count
Correct
7
Partially correct
2
Incorrect
3
What the Model Did Well… See the full description on the dataset page: https://huggingface.co/datasets/Acesif/blind-spot-fatima-institute-qwen3.5-0.8b.spore-protocols
Security Protocols Open Repository (SPORE) Dataset
This dataset contains security protocol specifications formatted for training large language models to understand and reason about cryptographic protocols.
Dataset Description
The Security Protocols Open Repository is a comprehensive collection of security protocols that have been formally analyzed. Each protocol specification includes:
Principal declarations (participants in the protocol)
Cryptographic primitives (keys… See the full description on the dataset page: https://huggingface.co/datasets/dassarthak18/spore-protocols.qwen2.5-3b-civil-engineering-blind-spots
Blind Spots of Qwen2.5-3B on Civil Engineering Domain Knowledge
Model Tested
Qwen/Qwen2.5-3B — a 3-billion parameter base language model released by Alibaba Cloud in September 2024. This is the base (not instruction-tuned) variant, selected because it represents the raw pre-training knowledge without task-specific fine-tuning.
How the Model Was Loaded
The model was loaded in a Google Colab notebook using a free T4 GPU:
from transformers import… See the full description on the dataset page: https://huggingface.co/datasets/HammamAkrami/qwen2.5-3b-civil-engineering-blind-spots.youtu-llm-2b-base-blind-spots
Youtu-LLM-2B-Base Blind Spots Evaluation Dataset
This dataset contains 75 evaluation prompts used to analyze the failure modes of tencent/Youtu-LLM-2B-Base,
a 1.96B parameter dense base language model released on December 31, 2025. Each row includes the input prompt, the expected answer, and the model’s
generated output obtained during inference on a Google Colab T4 GPU.
The prompts span 13 broad categories including arithmetic, logic, multilingual generation, instruction following… See the full description on the dataset page: https://huggingface.co/datasets/k-imtz/youtu-llm-2b-base-blind-spots.Sports_25k
WithinUsAI/Sports_25k — Master Scholars Academics (25k)
This dataset is designed for academic-grade fine-tuning of LLMs on sports rules, sports science, and quantitative sports analytics with a Tiny-Recursive-Model-friendly structure.
What’s inside (25,000 examples)
Task mix (fixed):
7,000 Fact-check / verification items (wrapper=verify_true_false, truth_mode=verifiable_fact)
10,000 Self-contained quantitative reasoning items (wrapper=minimal_chain… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Sports_25k.qwen3.5-4b-base-blind-spots
Qwen3.5-4B-Base Blind Spot Dataset
Model Tested
Qwen/Qwen3.5-4B-Base — a 4B-parameter multimodal base model (pre-trained, not instruction-tuned) released March 2, 2026 by the Qwen team. Hybrid Gated DeltaNet + Sparse MoE architecture.
!pip install -q git+https://github.com/huggingface/transformers.git accelerate bitsandbytes torch
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
model_id = "Qwen/Qwen3.5-4B-Base"
bnb_config =… See the full description on the dataset page: https://huggingface.co/datasets/Similoluwa/qwen3.5-4b-base-blind-spots.gemma-2-2b-blind-spots
Gemma-2-2B Base Model Blind Spots Dataset
Dataset Description
This dataset contains 10 carefully curated examples that highlight specific blind spots and failure modes of the google/gemma-2-2b base model. The examples span diverse categories of reasoning and computation where the base model demonstrates systematic weaknesses.
Model Tested: google/gemma-2-2b
Type: Base model (pre-trained, not instruction-tuned)
Parameters: 2.6B
Release Date: 2024
Methodology… See the full description on the dataset page: https://huggingface.co/datasets/SumaiyaMifra/gemma-2-2b-blind-spots.spotless-customer-service-training
Spotless Bin Co Customer Service Training Data
Training data for a customer service AI model for Spotless Bin Co, a residential trash can cleaning service.
Dataset Description
This dataset contains 8,776 conversational examples across 5 categories:
Category
Count
Description
FAQs
1,951
Frequently asked questions
Service
1,925
Service explanation dialogues
Objections
1,925
Objection handling examples
Booking
1,975
Booking flow conversations
Brand
1,000… See the full description on the dataset page: https://huggingface.co/datasets/rileyseaburg/spotless-customer-service-training.ayrton-1-qa-v2
Ayrton-1 QA v2
A question-answering dataset for supervised fine-tuning of language models on Formula 1 knowledge, 1950–2025.
This is the training corpus behind machina-sports/ayrton-1.
What's in it
Template-generated question/answer pairs covering:
Race results, championship standings, driver/constructor history (1950–2025, backed by Jolpica-F1).
Session-level telemetry and strategy — lap times, pit stops, stints, compound usage, top speeds (2018–2025, backed by FastF1).… See the full description on the dataset page: https://huggingface.co/datasets/machina-sports/ayrton-1-qa-v2.porqsport-indian-sportsbook
PorqSport Indian Sportsbook Knowledge
Curated Q&A knowledge for the /porqsport sportsbook agent.
Covers Indian exchange standards: Khai/Lagai, fancy & session markets, odds calculations,
UI buttons (back/lay, bet slip, MIN/MAX), market status signs, INR/UPI, STOMP feeds, and
guru365/robuzz industry API patterns.
Rows: 209
Usage
from datasets import load_dataset
ds = load_dataset("lavmauryaa/porqsport-indian-sportsbook", split="train")
Categories… See the full description on the dataset page: https://huggingface.co/datasets/lavmauryaa/porqsport-indian-sportsbook.nanbeige-3b-blind-spots
Blind Spots Evaluation: Nanbeige/Nanbeige4-3B-Base
Model Tested
Model name: Nanbeige4-3B-Base
Parameter count: 3B
Architecture: LlamaForCausalLM
Release date: 06 December 2025
Confirmation: This is a pure base model with no chat template applied. It requires manual completion or few-shot prompting for structured tasks.
How to Load the Model
Include this exact working code:
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch… See the full description on the dataset page: https://huggingface.co/datasets/afalaudn/nanbeige-3b-blind-spots.qwen3-1.7b-blind-spots
Blind Spots of a Frontier Base Model: Failure Analysis of Qwen3-1.7B-Base
Dataset Link
Public Hugging Face Dataset
1. Model Selection
To complete this challenge, I browsed recently released open models on
Hugging Face within the 0.6B--6B parameter range.
The model selected for analysis:
Model Name: Qwen3-1.7B-Base Parameter Size: 1.7B Modality: Text (Causal Language Model) Type: Base model (not instruction-tuned) Availability: Public on Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/Habibgm/qwen3-1.7b-blind-spots.spoiler_generation
Datasets used for spoiler generation task
Dataset Description
This dataset contains multiple datasets used for spoiler generation task for Clickbait spoiling competition for
training models based on Question Answering, text generation, or learn-to-rank problems.
Dataset Structure
This dataset has 5 main directories:
clf_data - this dataset was used to train a classifier for two generated spoilers to decide
which spoiler better matches a clickbait post. This… See the full description on the dataset page: https://huggingface.co/datasets/MateuszW/spoiler_generation.spontanous-speech-qa
Spontanous speech QA
This dataset contains QA pairs from the spontaneous speech subsection of the Danish Gigaword.
The dataset is created from the DDSC dataset and
filtered to only include QA pairs where the question is less than 20 tokens and the answer is
at least 4 tokens long.
To find out more about the creation see the accompanying script.
Blind_Spots_Dataset_SmolLM3-3B-Base
SmolLM3-3B-Base Blind Spots Dataset
Dataset Summary
This dataset documents 13 failure cases of the HuggingFaceTB/SmolLM3-3B-Base model, a 3-billion parameter base language model pretrained on 11.2 trillion tokens. Each row contains a text completion prompt, the expected correct output, the model's actual output, and the category. The dataset spans 13 distinct categories in which the SmolLM3-3B-Base model fails to work as expected.
Model Tested
Model:… See the full description on the dataset page: https://huggingface.co/datasets/Yanmife/Blind_Spots_Dataset_SmolLM3-3B-Base.tiny-aya-base-blind-spots
Blind Spots of a Frontier Base Model: Evaluation Dataset
This dataset documents blind spots discovered in a frontier open-weight base model through 19 structured evaluation tests. It was assembled as part of an assignment on identifying model weaknesses using the HelloBench evaluation framework.
Model Tested
CohereLabs/tiny-aya-base
Architecture: Transformer with Sliding Window Attention (SWA) (window size 4096, with RoPE) on three layers + one global attention layer… See the full description on the dataset page: https://huggingface.co/datasets/Mawube/tiny-aya-base-blind-spots.Blind-Spots-of-Frontier-Models_Multihop-blindspots-qwen35
Qwen3.5-2B-Base — Multi-Hop Reasoning Blind Spots
An evaluation dataset probing 18 Knowledge Graph-style reasoning tasks on
Qwen/Qwen3.5-2B-Base, tested in its
raw base (pre-training) form with no external graph attached. The dataset covers
parametric memory (probes 1–10, no passage provided), standard grounded reasoning
(probes 11–15, source passage included), and advanced grounded reasoning
(probes 16–18, passage provided but requiring implicit inference or contradiction… See the full description on the dataset page: https://huggingface.co/datasets/chayma-rhaiem/Blind-Spots-of-Frontier-Models_Multihop-blindspots-qwen35.qwen-1.5b-blind-spots
Qwen2.5-1.5B Blind Spots Dataset
Dataset Description
This dataset documents failure cases ("blind spots") of the base model
Qwen/Qwen2.5-1.5B — a 1.5-billion
parameter open-source language model released by Alibaba's Qwen team.
Each row contains a prompt, the answer a correct reasoner would give, and the
actual output the model produced — illustrating where it goes wrong.
Model Tested
Field
Value
Model
Qwen/Qwen2.5-1.5B
Parameters
1.5 B
Type… See the full description on the dataset page: https://huggingface.co/datasets/bozahbe21/qwen-1.5b-blind-spots.va-spoken-qa-agentvoice-stage2
Stage-2 spoken-QA training corpus (gemma4_talker)
97,697 train / 935 val spoken QA pairs. Questions: REAL user audio from
VoiceAssistant-400K (audio_q.*.tar, filenames match input_audio basenames in the
manifests). Answers: synthesized in ONE fixed agent voice (LibriSpeech
train-clean-100 narrator ref via ResembleAI/Chatterbox; audio_a.*.tar matching
assistant_audio). Manifests carry transcripts, answer text, and GLM-4-Voice speech
tokens for both sides (glm_in_tokens question /… See the full description on the dataset page: https://huggingface.co/datasets/z050209/va-spoken-qa-agentvoice-stage2.
