datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LiveSports-3K
LiveSports-3K Benchmark
News
[2025.05.12] We released the ASR transcripts for the CC track. See LiveSports-3K-CC.json for details.
Overview
LiveSports‑3K is a comprehensive benchmark for evaluating streaming video understanding capabilities of large language
and multimodal models. It consists of two evaluation tracks:
Closed Captions (CC) Track: Measures models’ ability to generate real‑time commentary aligned with the
ground‑truth ASR transcripts.
Question… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/LiveSports-3K.live-facts-snapshot
Live Facts Snapshot
A daily snapshot of verifiable, post-training-cutoff world-state facts — the kind of
ground truth language models cannot know from training data — exported through
Dynamic Feed, a live, verifiable data API whose every response
is Ed25519-signed. One file per day (data/YYYY-MM-DD.jsonl), one fact per line, and
every row carries its own source, source_url and measured_at.
Facts covered per day:
tool
facts
upstream source
licence
software_version… See the full description on the dataset page: https://huggingface.co/datasets/dynamicfeed/live-facts-snapshot.LiveClin
[ICLR'26] LiveClin: A Live Clinical Benchmark
📃 Paper •
🤗 Dataset •
💻 Code
LiveClin is a contamination-free, biannually updated clinical benchmark for evaluating large vision-language models on realistic, multi-stage clinical case reasoning with medical images and tables.
Each case presents a clinical scenario followed by a sequence of multiple-choice questions (MCQs) that mirror the progressive diagnostic workflow a clinician would follow — from initial… See the full description on the dataset page: https://huggingface.co/datasets/AQ-MedAI/LiveClin.liveqaThis is LiveQA, a Chinese dataset constructed from play-by-play live broadcast.
It contains 117k multiple-choice questions written by human commentators for over 1,670 NBA games,
which are collected from the Chinese Hupu website.liveqa_medical_trec2017
Dataset Card for LiveQA Medical from TREC 2017
The LiveQA'17 medical task focuses on consumer health question answering. Consumer health questions were received by the U.S. National Library of Medicine (NLM).
The dataset consists of constructed medical question-answer pairs for training and testing, with additional annotations that can be used to develop question analysis and question answering systems.
Please refer to our overview paper for more information about the constructed… See the full description on the dataset page: https://huggingface.co/datasets/hyesunyun/liveqa_medical_trec2017.LiveDRBench
Dataset Card for LiveDRBench: Deep Research as Claim Discovery
Arxiv Paper | Hugging Face Dataset | Evaluation Code
We propose a formal characterization of the deep research (DR) problem and introduce a new benchmark, LiveDRBench, to evaluate the performance of DR systems. To enable objective evaluation, we define DR using an intermediate output representation that encodes key claims uncovered during search—separating the reasoning challenge from surface-level report generation.… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/LiveDRBench.livek12bench
LiveK12Bench
A bilingual (Chinese/English) K12 high school benchmark dataset covering four STEM subjects: Mathematics, Physics, Chemistry, and Biology.
Questions are sourced from real Chinese high school exam papers and are available in both Chinese and English.
Code & Evaluation Framework: The full pipeline for exam parsing, model evaluation, and metrics is available at 👉 https://github.com/QQ-MM/LiveK12Bench
Dataset Structure
The dataset has four splits:… See the full description on the dataset page: https://huggingface.co/datasets/Shawn-wxh/livek12bench.LiveMathBench
Dataset Card for "LiveMathBench"
Homepage: https://open-compass.github.io/GPassK/
Repository: https://github.com/open-compass/GPassK
Paper: Are Your LLMs Capable of Stable Reasoning?
Introduction
LiveMathBench is a mathematical dataset, specifically designed to include challenging latest question sets from various mathematical competitions, aiming to avoid data contamination issues in existing LLMs and public math benchmarks.
Leaderboard
The Latest… See the full description on the dataset page: https://huggingface.co/datasets/opencompass/LiveMathBench.liveqa_trec2017
Dataset Card for LiveQA Medical from TREC 2017
The LiveQA'17 medical task focuses on consumer health question answering. Consumer health questions were received by the U.S. National Library of Medicine (NLM).
The dataset consists of constructed medical question-answer pairs for training and testing, with additional annotations that can be used to develop question analysis and question answering systems.
Please refer to our overview paper for more information about the constructed… See the full description on the dataset page: https://huggingface.co/datasets/katielink/liveqa_trec2017.Benchmark
Description
The document describes the LiveRAG benchmark.
For more details regarding Q&A generation see [1,2].
The LiveRAG benchmark includes 895 questions:
500 questions from Session 1, 500 questions from Session 2, with 105 shared questions from both Sessions
A total of 895 unique questions
Benchmark Fields
Field name
Description
Type
Remarks
Index
Benchmark index
int64 [0,1,...,894]
Question
DataMorgana question
String
Answer
DataMorgana ground… See the full description on the dataset page: https://huggingface.co/datasets/LiveRAG/Benchmark.LiveVQA-2025Wiki_Live_Challenge
Wiki Live Challenge Dataset
[English | 中文]
English
📖 Dataset Overview
This is the official dataset accompanying the Wiki Live Challenge benchmark. It contains Wikipedia Good Articles (GAs) as ground truth and research articles generated by leading deep research AI systems.
Wiki Live Challenge is the first live benchmark for evaluating Deep Research Agents (DRAs) on their ability to generate Wikipedia-quality articles. Unlike static benchmarks, Wiki Live… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/Wiki_Live_Challenge.livemath-v7-2603-2606
LiveMathematicianBench v7 — arXiv 2026-03 … 2026-06
An automated, refreshable benchmark of research-level mathematics multiple-choice questions,
generated from newly published arXiv math papers (contamination-resistant by construction).
Columbia University & Microsoft Research.
Each question is grounded in a theorem from a recent paper; distractors are adversarially crafted
from the proof sketch, and a multi-stage hardness pipeline keeps the final set difficult for frontier… See the full description on the dataset page: https://huggingface.co/datasets/hendrydong/livemath-v7-2603-2606.LiveCodeBench-CodeGenerationLiveProteinBench
LiveProteinBench
LiveProteinBench is a protein-centric benchmark with multiple-choice QA files
and rendered protein structure images for evaluating language and multimodal
models on protein understanding tasks.
License
The dataset content is released under the Creative Commons Attribution 4.0
International License (CC BY 4.0). See LICENSE for details.
Please also respect the applicable terms and citation requirements of any
upstream biological resources used to construct… See the full description on the dataset page: https://huggingface.co/datasets/LiveProteinBench/LiveProteinBench.browserbench-live-expanded
BrowserBench Live Expanded
Live-browser tasks derived only from the starting_url field of
Halluminate/BrowserBench at revision
aa56ce5e6331425c29878037fa8c169507deecdc.
Each source page was revisited live at 800×600. Up to three tasks were generated
from the current page observation: information, navigation, and interaction.
Historical prompts, results, screenshots, and ground-truth URLs were not used.
The canonical dataset contains all 876 task slots from 292 source pages. A… See the full description on the dataset page: https://huggingface.co/datasets/merve/browserbench-live-expanded.live-face-swap-support-qa
LiveFaceSwap AI Public Support Q&A
This dataset contains English question and answer pairs from the public LiveFaceSwap AI browser, desktop, and pricing FAQs, captured on 2026-09-17. Each row records its source page. The official website is the current source for product behavior and pricing; this snapshot can become outdated.
It is suitable for evaluating or prototyping retrieval over LiveFaceSwap AI product support content. It is not a face image dataset, a face swap training… See the full description on the dataset page: https://huggingface.co/datasets/LiveFaceSwapAI/live-face-swap-support-qa.livenewsbench-search-arms
LiveNewsBench Search Arms
Paired measurements of four language models answering the same 1,329 news
questions under four retrieval conditions. Every question was run in every
condition, so each row pairs with 13 others on task_key.
The release answers one question: how much of an agent's answer quality comes
from the model, and how much from the search system wrapped around it.
This is a derivative evaluation-results dataset, not the original
LiveNewsBench benchmark. The… See the full description on the dataset page: https://huggingface.co/datasets/BraintrustDataDev/livenewsbench-search-arms.LiveCodeBench-EvoSyn
EvoSyn-LiveCodeBench: Evolutionary Synthesized Coding Problems
Dataset Description
This dataset contains 231 high-quality coding problems synthesized and filtered using the EvoSyn framework.
Each problem includes diverse and reliable unit tests, specifically designed for reinforcement learning with verifiable rewards (RLVR).
Data Fields
We've adapted the original LiveCodeBench dataset structure, placing all unit tests into the public_test_cases field. This… See the full description on the dataset page: https://huggingface.co/datasets/Elynden/LiveCodeBench-EvoSyn.uk-live-music-glossary
UK Live Music Industry Glossary
The canonical structured glossary of UK live music industry terminology, published by GigXchange under CC BY 4.0.
Overview
Metric
Value
Terms defined
125
QA pairs
143
Role tags
4 (artist, venue, agent, promoter)
Categories
15
Cross-references
348 inter-term links
Language
British English (en-GB)
Domain
UK live music booking, performance, licensing, contracts
Splits
definitions… See the full description on the dataset page: https://huggingface.co/datasets/gigxchange/uk-live-music-glossary.uk-live-music-blog-corpus
UK Live Music Blog & Guides Corpus
109 long-form articles on the UK live music industry, published by GigXchange under CC BY 4.0. Written by working musicians and venue operators — not a content farm.
Overview
Metric
Value
Articles
109
Total words
305,841
Avg words/article
2,806
FAQ pairs
761
Topics
9
Date range
2026-03-01 to 2026-08-09
Language
British English (en-GB)
Domain
UK live music booking, fees, contracts, venues, city scenes… See the full description on the dataset page: https://huggingface.co/datasets/gigxchange/uk-live-music-blog-corpus.LiveVQA-Research-Preview
LIVEVQA: Live Visual Knowledge Seeking
Dataset Description
LIVEVQA is a benchmark dataset designed to evaluate the capabilities of Multimodal Large Language Models (MLLMs) in understanding and reasoning about live visual knowledge. Sourced from recent news articles (collected between March 14 and March 23, 2025), the dataset challenges models with questions requiring up-to-date, real-world information derived from images and associated news context.
The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/ONE-Lab/LiveVQA-Research-Preview.
