CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01livecodebench /code_generation LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code 🏠 Home Page • 💻 GitHub Repository • 🏆 Leaderboard • LiveCodeBench is a "live" updating benchmark for holistically evaluating code related capabilities of LLMs. Particularly, it evaluates LLMs across a range of capabilties including code generation, self-repair, test output prediction, and code execution. This is the code generation scenario of LiveCodeBench. It is also… See the full description on the dataset page: https://huggingface.co/datasets/livecodebench/code_generation.textn<1K32 likes5.3k downloads2y agoHugging Face02stdKonjac /LiveSports-3K LiveSports-3K Benchmark News [2025.05.12] We released the ASR transcripts for the CC track. See LiveSports-3K-CC.json for details. Overview LiveSports‑3K is a comprehensive benchmark for evaluating streaming video understanding capabilities of large language and multimodal models. It consists of two evaluation tracks: Closed Captions (CC) Track: Measures models’ ability to generate real‑time commentary aligned with the ground‑truth ASR transcripts. Question… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/LiveSports-3K.tabularvideo-text-to-text1K<n<10K5 likes1.5k downloads1y agoHugging Face03LiveMathematicianBench /LiveMathematicianBenchtextn<1K6 likes920 downloads2mo agoHugging Face04dynamicfeed /live-facts-snapshot Live Facts Snapshot A daily snapshot of verifiable, post-training-cutoff world-state facts — the kind of ground truth language models cannot know from training data — exported through Dynamic Feed, a live, verifiable data API whose every response is Ed25519-signed. One file per day (data/YYYY-MM-DD.jsonl), one fact per line, and every row carries its own source, source_url and measured_at. Facts covered per day: tool facts upstream source licence software_version… See the full description on the dataset page: https://huggingface.co/datasets/dynamicfeed/live-facts-snapshot.textquestion-answering1K<n<10K0 likes792 downloads20h agoHugging Face05birdsql /livesqlbench-base-lite-sqlite 🚀 LiveSQLBench-Base-Lite A dynamic, contamination‑free benchmark for evaluating LLMs on complex, real‑world text‑to‑SQL tasks. 🌐 LiveSQLBench Website • 🌐 BIRD-INTERACT Project Page • 📄 Paper • 💻 LiveSQLBench GitHub • 💻 BIRD-INTERACT GitHub Maintained by the 🦜 BIRD Team @ HKU & ☁️ Google Cloud 📊 LiveSQLBench Overview LiveSQLBench (BIRD-SQL Pro v0.5) is a contamination-free, continuously evolving benchmark designed to evaluate LLMs on complex, real-world… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/livesqlbench-base-lite-sqlite.texttable-question-answeringn<1K4 likes730 downloads4mo agoHugging Face06Live12 /btcv Beyond the Cranial Vault Dataset Dataset Description The Beyond the Cranial Vault dataset for multi-organ abdominal CT segmentation. This dataset contains CT scans with dense segmentation annotations. Dataset Details Modality: CT Target: 13 abdominal organs Format: NIfTI (.nii.gz) Dataset Structure Each sample in the JSONL file contains: { "image": "path/to/image.nii.gz", "mask": "path/to/mask.nii.gz", "label": ["organ1", "organ2", ...]… See the full description on the dataset page: https://huggingface.co/datasets/Live12/btcv.textimage-segmentationn<1K0 likes664 downloads7mo agoHugging Face07MedOtter /msd-liver Medical Segmentation Decathlon: Liver Dataset Description This is the Liver dataset from the Medical Segmentation Decathlon (MSD) challenge. The dataset contains CT scans with segmentation annotations for liver and liver tumor segmentation. Dataset Details Modality: CT Task: Task03_Liver Target: liver and liver tumors Format: NIfTI (.nii.gz) Dataset Structure Each sample in the JSONL file contains: { "image":… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/msd-liver.textimage-segmentationn<1K0 likes599 downloads11mo agoHugging Face08nvidia /LiveCodeBench-CPP LiveCodeBench-CPP: An Extension of LiveCodeBench for Contamination Free Evaluation in C++ Overview LiveCodeBench-CPP includes 454 problems from the release_v6 of LiveCodeBench, covering the period from October 2024 to May 2025. These problems are sourced from AtCoder (287 problems) and LeetCode (167 problems). AtCoder Problems: These require generated solutions to read inputs from standard input (stdin) and write outputs to standard output (stdout). For unit testing, the… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/LiveCodeBench-CPP.textn<1K4 likes553 downloads1y agoHugging Face09birdsql /livesqlbench-base-full-v1 🚀 LiveSQLBench-Base-Full-v1 A dynamic, contamination‑free benchmark for evaluating LLMs on complex, real‑world text‑to‑SQL tasks. 🌐 Website/Leaderboard • 📄 Paper (coming soon) • 💻 GitHub • 🗄️ LiveSQLBench-Base-Lite • 🗄️ LiveSQLBench-Large-v1 • 🗄️ Bird-Interact (ICLR 2026 Oral) Maintained by the 🦜 BIRD Team @ HKU & ☁️ Google Cloud 📊 LiveSQLBench Overview LiveSQLBench (BIRD-SQL Pro v0.5) is a contamination-free, continuously evolving benchmark designed to… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/livesqlbench-base-full-v1.textn<1K3 likes527 downloads4mo agoHugging Face10bzantium /livecodebench LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code Note: This is a clone of livecodebench/code_generation_lite updated to work with recent versions of the datasets library. The original repository uses a Python loading script which is no longer supported. This version provides the same data using the standard JSONL format for compatibility. Dataset Description LiveCodeBench is a "live" updating benchmark for holistically… See the full description on the dataset page: https://huggingface.co/datasets/bzantium/livecodebench.text1K<n<10K0 likes444 downloads10mo agoHugging Face11birdsql /livesqlbench-large-v1 🚀 LiveSQLBench-Large-v1 A dynamic, contamination‑free benchmark for evaluating LLMs on complex, real‑world text‑to‑SQL tasks at industrial scale. 🌐 Website/Leaderboard • 📄 Paper (coming soon) • 💻 GitHub • 🗄️ LiveSQLBench-Base-Lite • 🗄️ LiveSQLBench-Base-Full-v1 • 🗄️ Bird-Interact (ICLR 2026 Oral) Maintained by the 🦜 BIRD Team @ HKU & ☁️ Google Cloud 📊 LiveSQLBench Overview LiveSQLBench (BIRD-SQL Pro v0.5) is a contamination-free, continuously evolving… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/livesqlbench-large-v1.textn<1K4 likes407 downloads7mo agoHugging Face12ICIP /LiveMCPBench LiveMCPBench: Can Agents Navigate an Ocean of MCP Tools? Benchmarking the agent in real-world tasks within a large-scale MCP toolset. 🌐 Website   |   📄 Paper   |   💻 Code   |   🏆 Leaderboard   |   🙏 Citation Dataset Description LiveMCPBench is the first comprehensive benchmark designed to evaluate LLM agents at scale across diverse Model Context Protocol (MCP) servers. It comprises 95 real-world tasks grounded in the MCP ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/ICIP/LiveMCPBench.audioimage-text-to-textn<1K9 likes374 downloads1y agoHugging Face13AQ-MedAI /LiveClin [ICLR'26] LiveClin: A Live Clinical Benchmark 📃 Paper • 🤗 Dataset • 💻 Code LiveClin is a contamination-free, biannually updated clinical benchmark for evaluating large vision-language models on realistic, multi-stage clinical case reasoning with medical images and tables. Each case presents a clinical scenario followed by a sequence of multiple-choice questions (MCQs) that mirror the progressive diagnostic workflow a clinician would follow — from initial… See the full description on the dataset page: https://huggingface.co/datasets/AQ-MedAI/LiveClin.imagequestion-answering1K<n<10K9 likes348 downloads7mo agoHugging Face14birdsql /livesqlbench-base-lite 🚀 LiveSQLBench-Base-Lite A dynamic, contamination‑free benchmark for evaluating LLMs on complex, real‑world text‑to‑SQL tasks. 🌐 Website • 📄 Paper (coming soon) • 💻 GitHub • 🗄️ LiveSQLBench-Base-Full-v1 • 🗄️ LiveSQLBench-Large-v1 • 🗄️ Bird-Interact (ICLR 2026 Oral) Maintained by the 🦜 BIRD Team @ HKU & ☁️ Google Cloud 📊 LiveSQLBench Overview LiveSQLBench (BIRD-SQL Pro v0.5) is a contamination-free, continuously evolving benchmark designed to evaluate LLMs on… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/livesqlbench-base-lite.textn<1K5 likes322 downloads7mo agoHugging Face15iMeanAI /Mind2Web-LiveDataset: Mind2Web-Live Github: https://github.com/iMeanAI/WebCanvas textn<1K16 likes268 downloads2y agoHugging Face16Forival /LiveBrowseComp LiveBrowseComp Dataset Basic Information Filename: LiveBrowseComp.jsonl Number of questions: 335 IDs: 0-334 (continuous) Format: JSONL, with one JSON object per line Field Description Field Type Description idx int Question ID, continuous from 0 to 334 problem str Problem description, a complex reasoning question composed of multiple clues answer str Reference answer Question Characteristics The dataset covers… See the full description on the dataset page: https://huggingface.co/datasets/Forival/LiveBrowseComp.textn<1K1 likes231 downloads4mo agoHugging Face17hyesunyun /liveqa_medical_trec2017 Dataset Card for LiveQA Medical from TREC 2017 The LiveQA'17 medical task focuses on consumer health question answering. Consumer health questions were received by the U.S. National Library of Medicine (NLM). The dataset consists of constructed medical question-answer pairs for training and testing, with additional annotations that can be used to develop question analysis and question answering systems. Please refer to our overview paper for more information about the constructed… See the full description on the dataset page: https://huggingface.co/datasets/hyesunyun/liveqa_medical_trec2017.textquestion-answeringn<1K8 likes205 downloads3y agoHugging Face18livebench /liveswebench LiveSWEBench Tasks This dataset contains all task instances for the LiveSWEBench benchmark. Tasks are stored in the following format: repo_name (str): the name of the repository for this task task_num (int): the original PR number for the task, used now as an identifier gold_patch (str): the actual changes made in the PR to resolve the issue in this task test_patch (str): the changes made to test files to validate the task solution edit_patch (str): a subset of the gold patch… See the full description on the dataset page: https://huggingface.co/datasets/livebench/liveswebench.textn<1K3 likes205 downloads1y agoHugging Face19OpenClaw /clawhub-security-signals-live ClawHub Security Signals Live This dataset is the refreshed ClawHub security-signals corpus for scanner testing, prompt regression checks, and operational research against recent public ClawHub skills. It is a moving dataset, not the fixed paper benchmark. main is expected to change when the ClawHub security dataset snapshot workflow publishes a new sanitized export. Pin a Hugging Face revision or commit when you need reproducibility. For the frozen research-paper snapshot, use… See the full description on the dataset page: https://huggingface.co/datasets/OpenClaw/clawhub-security-signals-live.tabulartext-classification10K<n<100K0 likes205 downloads5d agoHugging Face20livebench /liveswebench-patchestextn<1K1 likes193 downloads1y agoHugging Face21codezakh /dataenvgym-livecodebench-solutionstextn<1K0 likes188 downloads2y agoHugging Face22stanfordnlp /nnetnav-livetext10K<n<100K1 likes181 downloads2y agoHugging Face23oaaoaa /LiveGamingBenchmarkgatedtabular1K<n<10K3 likes106 downloads4mo agoHugging Face24justicedao /federal-register-live-graphrag-research-20260810 Federal Register live GraphRAG (research) Local LCR-071 live pipeline output for the 2026-08-10 cutoff (11,784 documents, CUDA thenlper/gte-small). This Hub copy is a research snapshot. It is not a current-bundle and does not replace justicedao/ipfs_federal_register. LCR-084 remains open. Official Federal Register publications remain the authority. Hub git directories may contain at most 10,000 files. Document bodies beyond that cap are stored under corpus/bodies-part2/ rather… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/federal-register-live-graphrag-research-20260810.tabularn<1K0 likes105 downloads19d agoHugging Face25ruiyang-medinfo /GlobMed_LiveQA 🌍 GlobMed: LiveQA GlobMed_LiveQA covers 20 languages, including 13 high-resource languages (Arabic, Chinese, English, French, German, Hindi, Indonesian, Japanese, Korean, Portuguese, Russian, Spanish, and Thai) and 7 low-resource languages (Bengali, Malay, Swahili, Urdu, Wolof, Yoruba, and Zulu). Code ar bn zh en fr de hi id ja ko ms pt ru es sw th ur woyo zu Language Arabic Bengali Chinese English French German Hindi Indonesian Japanese Korean Malay Portuguese Russian… See the full description on the dataset page: https://huggingface.co/datasets/ruiyang-medinfo/GlobMed_LiveQA.text10K<n<100K0 likes102 downloads8mo agoHugging Face26fm-universe /Live-FM-Bench Introduction This dataset Live-FM-bench is continuously updated and contamination-free evaluation benchmark of LLMs for program verification (a.k.a., formal specification generation). Currently, it contains 360 C programs under verification together with the properties to be verified. This dataset can be used for: Specification generation task (Code2Proof): given program and properties to be verified as input, output program with specification that can pass the prover.… See the full description on the dataset page: https://huggingface.co/datasets/fm-universe/Live-FM-Bench.textn<1K2 likes100 downloads2mo agoHugging Face27opencompass /LiveMathBenchgated Dataset Card for "LiveMathBench" Homepage: https://open-compass.github.io/GPassK/ Repository: https://github.com/open-compass/GPassK Paper: Are Your LLMs Capable of Stable Reasoning? Introduction LiveMathBench is a mathematical dataset, specifically designed to include challenging latest question sets from various mathematical competitions, aiming to avoid data contamination issues in existing LLMs and public math benchmarks. Leaderboard The Latest… See the full description on the dataset page: https://huggingface.co/datasets/opencompass/LiveMathBench.textquestion-answeringn<1K13 likes99 downloads1y agoHugging Face28BenchEvolver /livecodebench-plus LiveCodeBench-v6-Plus A curated coding benchmark of 91 problems selected by hardness/discrimination (lcb-v6-plus). It combines two sources, all in one clean schema: 64 evolved problems — mutated/evolved variants from LiveCodeBench-v6 (each carries its seed_problem). 27 original problems — un-evolved AtCoder problems taken directly from livecodebench/code_generation_lite release v6 (seed_problem is null). About BenchEvolver The evolved problems were produced by… See the full description on the dataset page: https://huggingface.co/datasets/BenchEvolver/livecodebench-plus.texttext-generationn<1K0 likes98 downloads4mo agoHugging Face29bzantium /ko-livecodebench Ko-LiveCodeBench This dataset is the livecodebench/code_generation_lite dataset with the question_content field translated to Korean. Dataset Versions The dataset provides multiple configurations (subsets) corresponding to different release versions: release_v1: Problems released between May 2023 and Mar 2024 (400 problems) release_v2: Problems released between May 2023 and May 2024 (511 problems) release_v3: Problems released between May 2023 and Jul 2024 (612 problems)… See the full description on the dataset page: https://huggingface.co/datasets/bzantium/ko-livecodebench.text1K<n<10K1 likes91 downloads10mo agoHugging Face30joeygambino /twitch-top-live-streams-metadata Top Twitch Channels Live Viewership & Broadcast Metadata Overview This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines. Source Actor: captainhandsome/twitch-live-streams-scraper Dataset Page: Public sample and schema Preconfigured Run Task: captainhandsome/twitch-live-fortnite-streams… See the full description on the dataset page: https://huggingface.co/datasets/joeygambino/twitch-top-live-streams-metadata.imageothern<1K0 likes81 downloads8d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.