datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_generation
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
🏠 Home Page •
💻 GitHub Repository •
🏆 Leaderboard •
LiveCodeBench is a "live" updating benchmark for holistically evaluating code related capabilities of LLMs.
Particularly, it evaluates LLMs across a range of capabilties including code generation, self-repair, test output prediction, and code execution.
This is the code generation scenario of LiveCodeBench. It is also… See the full description on the dataset page: https://huggingface.co/datasets/livecodebench/code_generation.LiveSports-3K
LiveSports-3K Benchmark
News
[2025.05.12] We released the ASR transcripts for the CC track. See LiveSports-3K-CC.json for details.
Overview
LiveSports‑3K is a comprehensive benchmark for evaluating streaming video understanding capabilities of large language
and multimodal models. It consists of two evaluation tracks:
Closed Captions (CC) Track: Measures models’ ability to generate real‑time commentary aligned with the
ground‑truth ASR transcripts.
Question… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/LiveSports-3K.LiveMathematicianBenchlive-facts-snapshot
Live Facts Snapshot
A daily snapshot of verifiable, post-training-cutoff world-state facts — the kind of
ground truth language models cannot know from training data — exported through
Dynamic Feed, a live, verifiable data API whose every response
is Ed25519-signed. One file per day (data/YYYY-MM-DD.jsonl), one fact per line, and
every row carries its own source, source_url and measured_at.
Facts covered per day:
tool
facts
upstream source
licence
software_version… See the full description on the dataset page: https://huggingface.co/datasets/dynamicfeed/live-facts-snapshot.livesqlbench-base-lite-sqlite
🚀 LiveSQLBench-Base-Lite
A dynamic, contamination‑free benchmark for evaluating LLMs on complex, real‑world text‑to‑SQL tasks.
🌐 LiveSQLBench Website • 🌐 BIRD-INTERACT Project Page • 📄 Paper • 💻 LiveSQLBench GitHub • 💻 BIRD-INTERACT GitHub
Maintained by the 🦜 BIRD Team @ HKU & ☁️ Google Cloud
📊 LiveSQLBench Overview
LiveSQLBench (BIRD-SQL Pro v0.5) is a contamination-free, continuously evolving benchmark designed to evaluate LLMs on complex, real-world… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/livesqlbench-base-lite-sqlite.btcv
Beyond the Cranial Vault Dataset
Dataset Description
The Beyond the Cranial Vault dataset for multi-organ abdominal CT segmentation. This dataset contains CT scans with dense segmentation annotations.
Dataset Details
Modality: CT
Target: 13 abdominal organs
Format: NIfTI (.nii.gz)
Dataset Structure
Each sample in the JSONL file contains:
{
"image": "path/to/image.nii.gz",
"mask": "path/to/mask.nii.gz",
"label": ["organ1", "organ2", ...]… See the full description on the dataset page: https://huggingface.co/datasets/Live12/btcv.msd-liver
Medical Segmentation Decathlon: Liver
Dataset Description
This is the Liver dataset from the Medical Segmentation Decathlon (MSD) challenge. The dataset contains CT scans with segmentation annotations for liver and liver tumor segmentation.
Dataset Details
Modality: CT
Task: Task03_Liver
Target: liver and liver tumors
Format: NIfTI (.nii.gz)
Dataset Structure
Each sample in the JSONL file contains:
{
"image":… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/msd-liver.LiveCodeBench-CPP
LiveCodeBench-CPP: An Extension of LiveCodeBench for Contamination Free Evaluation in C++
Overview
LiveCodeBench-CPP includes 454 problems from the release_v6 of LiveCodeBench, covering the period from October 2024 to May 2025. These problems are sourced from AtCoder (287 problems) and LeetCode (167 problems).
AtCoder Problems: These require generated solutions to read inputs from standard input (stdin) and write outputs to standard output (stdout). For unit testing, the… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/LiveCodeBench-CPP.livesqlbench-base-full-v1
🚀 LiveSQLBench-Base-Full-v1
A dynamic, contamination‑free benchmark for evaluating LLMs on complex, real‑world text‑to‑SQL tasks.
🌐 Website/Leaderboard • 📄 Paper (coming soon) • 💻 GitHub • 🗄️ LiveSQLBench-Base-Lite • 🗄️ LiveSQLBench-Large-v1 • 🗄️ Bird-Interact (ICLR 2026 Oral)
Maintained by the 🦜 BIRD Team @ HKU & ☁️ Google Cloud
📊 LiveSQLBench Overview
LiveSQLBench (BIRD-SQL Pro v0.5) is a contamination-free, continuously evolving benchmark designed to… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/livesqlbench-base-full-v1.livecodebench
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Note: This is a clone of livecodebench/code_generation_lite updated to work with recent versions of the datasets library. The original repository uses a Python loading script which is no longer supported. This version provides the same data using the standard JSONL format for compatibility.
Dataset Description
LiveCodeBench is a "live" updating benchmark for holistically… See the full description on the dataset page: https://huggingface.co/datasets/bzantium/livecodebench.livesqlbench-large-v1
🚀 LiveSQLBench-Large-v1
A dynamic, contamination‑free benchmark for evaluating LLMs on complex, real‑world text‑to‑SQL tasks at industrial scale.
🌐 Website/Leaderboard • 📄 Paper (coming soon) • 💻 GitHub • 🗄️ LiveSQLBench-Base-Lite • 🗄️ LiveSQLBench-Base-Full-v1 • 🗄️ Bird-Interact (ICLR 2026 Oral)
Maintained by the 🦜 BIRD Team @ HKU & ☁️ Google Cloud
📊 LiveSQLBench Overview
LiveSQLBench (BIRD-SQL Pro v0.5) is a contamination-free, continuously evolving… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/livesqlbench-large-v1.LiveMCPBench
LiveMCPBench: Can Agents Navigate an Ocean of MCP Tools?
Benchmarking the agent in real-world tasks within a large-scale MCP toolset.
🌐 Website |
📄 Paper |
💻 Code |
🏆 Leaderboard
|
🙏 Citation
Dataset Description
LiveMCPBench is the first comprehensive benchmark designed to evaluate LLM agents at scale across diverse Model Context Protocol (MCP) servers. It comprises 95 real-world tasks grounded in the MCP ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/ICIP/LiveMCPBench.LiveClin
[ICLR'26] LiveClin: A Live Clinical Benchmark
📃 Paper •
🤗 Dataset •
💻 Code
LiveClin is a contamination-free, biannually updated clinical benchmark for evaluating large vision-language models on realistic, multi-stage clinical case reasoning with medical images and tables.
Each case presents a clinical scenario followed by a sequence of multiple-choice questions (MCQs) that mirror the progressive diagnostic workflow a clinician would follow — from initial… See the full description on the dataset page: https://huggingface.co/datasets/AQ-MedAI/LiveClin.livesqlbench-base-lite
🚀 LiveSQLBench-Base-Lite
A dynamic, contamination‑free benchmark for evaluating LLMs on complex, real‑world text‑to‑SQL tasks.
🌐 Website • 📄 Paper (coming soon) • 💻 GitHub • 🗄️ LiveSQLBench-Base-Full-v1 • 🗄️ LiveSQLBench-Large-v1 • 🗄️ Bird-Interact (ICLR 2026 Oral)
Maintained by the 🦜 BIRD Team @ HKU & ☁️ Google Cloud
📊 LiveSQLBench Overview
LiveSQLBench (BIRD-SQL Pro v0.5) is a contamination-free, continuously evolving benchmark designed to evaluate LLMs on… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/livesqlbench-base-lite.Mind2Web-LiveDataset: Mind2Web-Live
Github: https://github.com/iMeanAI/WebCanvas
LiveBrowseComp
LiveBrowseComp Dataset
Basic Information
Filename: LiveBrowseComp.jsonl
Number of questions: 335
IDs: 0-334 (continuous)
Format: JSONL, with one JSON object per line
Field Description
Field
Type
Description
idx
int
Question ID, continuous from 0 to 334
problem
str
Problem description, a complex reasoning question composed of multiple clues
answer
str
Reference answer
Question Characteristics
The dataset covers… See the full description on the dataset page: https://huggingface.co/datasets/Forival/LiveBrowseComp.liveqa_medical_trec2017
Dataset Card for LiveQA Medical from TREC 2017
The LiveQA'17 medical task focuses on consumer health question answering. Consumer health questions were received by the U.S. National Library of Medicine (NLM).
The dataset consists of constructed medical question-answer pairs for training and testing, with additional annotations that can be used to develop question analysis and question answering systems.
Please refer to our overview paper for more information about the constructed… See the full description on the dataset page: https://huggingface.co/datasets/hyesunyun/liveqa_medical_trec2017.liveswebench
LiveSWEBench Tasks
This dataset contains all task instances for the LiveSWEBench benchmark. Tasks are stored in the following format:
repo_name (str): the name of the repository for this task
task_num (int): the original PR number for the task, used now as an identifier
gold_patch (str): the actual changes made in the PR to resolve the issue in this task
test_patch (str): the changes made to test files to validate the task solution
edit_patch (str): a subset of the gold patch… See the full description on the dataset page: https://huggingface.co/datasets/livebench/liveswebench.clawhub-security-signals-live
ClawHub Security Signals Live
This dataset is the refreshed ClawHub security-signals corpus for scanner testing, prompt regression checks, and operational research against recent public ClawHub skills.
It is a moving dataset, not the fixed paper benchmark. main is expected to change when the ClawHub security dataset snapshot workflow publishes a new sanitized export. Pin a Hugging Face revision or commit when you need reproducibility.
For the frozen research-paper snapshot, use… See the full description on the dataset page: https://huggingface.co/datasets/OpenClaw/clawhub-security-signals-live.liveswebench-patchesdataenvgym-livecodebench-solutionsnnetnav-liveLiveGamingBenchmarkfederal-register-live-graphrag-research-20260810
Federal Register live GraphRAG (research)
Local LCR-071 live pipeline output for the 2026-08-10 cutoff (11,784 documents,
CUDA thenlper/gte-small). This Hub copy is a research snapshot.
It is not a current-bundle and does not replace
justicedao/ipfs_federal_register. LCR-084 remains open. Official Federal
Register publications remain the authority.
Hub git directories may contain at most 10,000 files. Document bodies beyond
that cap are stored under corpus/bodies-part2/ rather… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/federal-register-live-graphrag-research-20260810.GlobMed_LiveQA
🌍 GlobMed: LiveQA
GlobMed_LiveQA covers 20 languages, including 13 high-resource languages (Arabic, Chinese, English, French, German, Hindi, Indonesian, Japanese, Korean, Portuguese, Russian, Spanish, and Thai) and 7 low-resource languages (Bengali, Malay, Swahili, Urdu, Wolof, Yoruba, and Zulu).
Code
ar
bn
zh
en
fr
de
hi
id
ja
ko
ms
pt
ru
es
sw
th
ur
woyo
zu
Language
Arabic
Bengali
Chinese
English
French
German
Hindi
Indonesian
Japanese
Korean
Malay
Portuguese
Russian… See the full description on the dataset page: https://huggingface.co/datasets/ruiyang-medinfo/GlobMed_LiveQA.Live-FM-Bench
Introduction
This dataset Live-FM-bench is continuously updated and contamination-free evaluation benchmark of LLMs for program verification (a.k.a., formal specification generation). Currently, it contains 360 C programs under verification together with the properties to be verified.
This dataset can be used for:
Specification generation task (Code2Proof): given program and properties to be verified as input, output program with specification that can pass the prover.… See the full description on the dataset page: https://huggingface.co/datasets/fm-universe/Live-FM-Bench.LiveMathBench
Dataset Card for "LiveMathBench"
Homepage: https://open-compass.github.io/GPassK/
Repository: https://github.com/open-compass/GPassK
Paper: Are Your LLMs Capable of Stable Reasoning?
Introduction
LiveMathBench is a mathematical dataset, specifically designed to include challenging latest question sets from various mathematical competitions, aiming to avoid data contamination issues in existing LLMs and public math benchmarks.
Leaderboard
The Latest… See the full description on the dataset page: https://huggingface.co/datasets/opencompass/LiveMathBench.livecodebench-plus
LiveCodeBench-v6-Plus
A curated coding benchmark of 91 problems selected by hardness/discrimination
(lcb-v6-plus). It combines two sources, all in one clean schema:
64 evolved problems — mutated/evolved variants from LiveCodeBench-v6
(each carries its seed_problem).
27 original problems — un-evolved AtCoder problems taken directly from
livecodebench/code_generation_lite
release v6 (seed_problem is null).
About BenchEvolver
The evolved problems were produced by… See the full description on the dataset page: https://huggingface.co/datasets/BenchEvolver/livecodebench-plus.ko-livecodebench
Ko-LiveCodeBench
This dataset is the livecodebench/code_generation_lite dataset with the question_content field translated to Korean.
Dataset Versions
The dataset provides multiple configurations (subsets) corresponding to different release versions:
release_v1: Problems released between May 2023 and Mar 2024 (400 problems)
release_v2: Problems released between May 2023 and May 2024 (511 problems)
release_v3: Problems released between May 2023 and Jul 2024 (612 problems)… See the full description on the dataset page: https://huggingface.co/datasets/bzantium/ko-livecodebench.twitch-top-live-streams-metadata
Top Twitch Channels Live Viewership & Broadcast Metadata
Overview
This dataset contains clean, structured public data exported directly from production runs of Apify actors.
It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines.
Source Actor: captainhandsome/twitch-live-streams-scraper
Dataset Page: Public sample and schema
Preconfigured Run Task: captainhandsome/twitch-live-fortnite-streams… See the full description on the dataset page: https://huggingface.co/datasets/joeygambino/twitch-top-live-streams-metadata.
