CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01livecodebench /code_generation_liteLiveCodeBench is a temporaly updating benchmark for code generation. Please check the homepage: https://livecodebench.github.io/.n<1K104 likes110k downloads1y agoHugging Face02princeton-nlp /SWE-bench_Lite Dataset Summary SWE-bench Lite is subset of SWE-bench, a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? Want to run inference now? This dataset only contains the… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_Lite.textn<1K66 likes98k downloads2y agoHugging Face03R2E-Gym /R2E-Gym-Litetabular10K<n<100K1 likes68k downloads2y agoHugging Face04SWE-bench /SWE-bench_Lite Dataset Summary SWE-bench Lite is subset of SWE-bench, a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? Want to run inference now? This dataset only contains the… See the full description on the dataset page: https://huggingface.co/datasets/SWE-bench/SWE-bench_Lite.textn<1K25 likes44k downloads1mo agoHugging Face05CohereLabs /Global-MMLU-Lite Releases: Version 3.0 (May 2026): GMMLU Lite 3.0 release with 5 new languages: Czech, Hungarian, Italian (updated), Oriya, Slovak and Tajik Version 2.0 (Dec 2025): GMMLU Lite 2.0 release with 3 new languages: Albanian, Burmese and Welsh Version 1.0 (Dec 2024): GMMLU Lite initial release with 15 languages. Dataset Summary Global-MMLU-Lite is a multilingual evaluation set spanning 23 languages, including English. It is "lite" version of the original Global-MMLU… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/Global-MMLU-Lite.text10K<n<100K43 likes30k downloads4mo agoHugging Face06SWE-Gym /SWE-Gym-LiteSWE-Gym Lite contains 230 instances sourced from 11 Python repos, following SWE-Bench Lite data collection procedure. Get started at project page github.com/SWE-Gym/SWE-Gym textn<1K3 likes13k downloads2y agoHugging Face07lmms-lab-encoder /LMMs-Eval-Liteimage1K<n<10K7 likes5.9k downloads2y agoHugging Face08initiacms /XLRS-Bench-lite 🐙GitHub Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench 📜Dataset License Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from: DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply). ITCVDLicensed under CC-BY-NC-SA-4.0. MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench-lite.imagevisual-question-answering1K<n<10K4 likes5.8k downloads11mo agoHugging Face09li-lab /MMLU-ProX-Lite MMLU-ProX-Lite MMLU-ProX is a multilingual benchmark that builds upon MMLU-Pro, extending to 29 typologically diverse languages, designed to evaluate large language models' reasoning capabilities across linguistic and cultural boundaries. Github | Paper News [2025/08] 🎉 MMLU-ProX was accepted by EMNLP 2025 Main Conference! [2025/05] MMLU-ProX now contains 29 languages, all available on Huggingface. [2025/03] MMLU-ProX is now available on Huggingface. [2025/03] We are… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/MMLU-ProX-Lite.tabular10K<n<100K3 likes5.1k downloads1y agoHugging Face10sailplane /SWE-bench_Lite_filtered0 likes4.7k downloads2y agoHugging Face11ystemsrx /Erotic_Literature_CollectionEnglish 中文色情文学数据集合集 概述 本仓库包含了51个中文色情文学数据集。每个数据集由短篇色情小说、个人色情经验及其他形式的色情内容组成。数据集的格式为JSON,每个文件包含一个对象数组,每个对象代表一篇文档: [ {"text": "document"}, {"text": "document"} ] 这些数据集可用于语言模型的预训练,经过适当调整后也可用于模型的微调。 数据集格式 文件格式: JSON 内容: 短篇色情小说、个人色情经验及其他色情内容 结构: 每个文件包含一个对象数组 每个对象包含一个键 "text",其值为相应的文档内容 使用方法 这些数据集主要用于研究目的,特别是在语言模型的开发和微调中使用。由于内容的敏感性,用户应谨慎处理这些数据集,并确保遵守当地的法律法规及相关指导原则。 示例用法 import json # 加载数据集with open('path_to_json_file.json', 'r'… See the full description on the dataset page: https://huggingface.co/datasets/ystemsrx/Erotic_Literature_Collection.texttext-generation10K<n<100K228 likes4.3k downloads2y agoHugging Face12cua-lite /lite.cuaworld-assets lite.cuaworld materials Maintained environment materials for cua-lite's lite.cuaworld.* software environments — forked from cmu-l3/gym-anything (CUA-World, MIT) and curated/edited/expanded by us. Published as the Hugging Face dataset cua-lite/lite.cuaworld-assets. This repo holds content only (no engine code): per-environment env.json, install/setup scripts/, per-task assets (tasks/<task>/…), data/, config/, assets/, optional post_build.sh, and a curated registered.json. The… See the full description on the dataset page: https://huggingface.co/datasets/cua-lite/lite.cuaworld-assets.0 likes4.1k downloads2mo agoHugging Face13LiteFold /PDB PDB mmCIF Entry Index The Protein Data Bank is the single global archive of experimentally-determined 3D structures of biological macromolecules, established in 1971 and now holding well over 230,000 entries. It stores atomic coordinates for proteins, nucleic acids, and their complexes determined by X-ray crystallography, cryo-EM, NMR, micro-electron diffraction, and integrative methods, along with the underlying experimental data (structure factors, EM maps, NMR restraints) and… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/PDB.tabular10K<n<100K2 likes3.9k downloads4mo agoHugging Face14lighteval /code_generation_lite LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code 🏠 Home Page • 💻 GitHub Repository • 🏆 Leaderboard • 📄 Paper Change Log Since LiveCodeBench is a continuously updated benchmark, we provide different versions of the dataset. Particularly, we provide the following versions of the dataset: release_v1: The initial release of the dataset with problems released between May 2023 and Mar 2024 containing 400… See the full description on the dataset page: https://huggingface.co/datasets/lighteval/code_generation_lite.text10K<n<100K5 likes3.8k downloads1y agoHugging Face15Workspace-Bench /Workspace-Bench-Lite Workspace-Bench-Lite A Lightweight Subset of Workspace-Bench for Fast and Cost-Efficient Evaluation Overview • LeaderBoard • Distribution • Quick Start • Changelog • Citation Overview Workspace-Bench-Lite is the Lite split of Workspace-Bench 1.0, designed for fast iteration and lower-cost benchmarking while preserving the core evaluation setting of the full benchmark. It contains 100 tasks selected from the full Workspace-Bench and is intended to… See the full description on the dataset page: https://huggingface.co/datasets/Workspace-Bench/Workspace-Bench-Lite.textn<1K2 likes3.8k downloads3mo agoHugging Face16sailplane /SWE-bench_Lite_filtered_10 likes3.3k downloads2y agoHugging Face17Chat-Error /book2-lite-cleanedtext10K<n<100K2 likes2.9k downloads3y agoHugging Face18Geralt-Targaryen /Literature-zhA composite of Chinese books, papers, legal documents, and patents from common crawl. Data cleaning Text with more than 2% of non-Latin, non-Chinese characters are removed. Text with large portions of special characters are removed. Traditional Chinese is converted to simplified Chinese. Model filtering Qwen2.5-32B-Instruct is used to generate language quality annotation (on a scale of 1-5) for 398K Chinese samples and 250K English samples. An XLM-RoBERT-large classifier is trained with… See the full description on the dataset page: https://huggingface.co/datasets/Geralt-Targaryen/Literature-zh.text10M<n<100M5 likes2.9k downloads1y agoHugging Face19yifishbossman /financial-analyst-data-lite financial-analyst-data-lite EN: A-share historical OHLCV + valuation + financials + TDX F10 events, packaged in Qlib binary + Parquet formats. Companion dataset for financial-analyst — a 14-agent single-stock deep-dive research workstation. 中文: A 股历史行情 + 估值 + 财报 + TDX F10 事件数据集, Qlib 二进制 + Parquet 双格式打包. 配套 financial-analyst — 14 Agent 个股深度研究工作站使用. Published / 发布: 2026-05-24 · Size / 体量: ~2.72 GB · License: Apache 2.0 📊 Three Preset Tiers / 三档预设 Pick the tier that… See the full description on the dataset page: https://huggingface.co/datasets/yifishbossman/financial-analyst-data-lite.texttime-series-forecasting1M<n<10M1 likes2.9k downloads4mo agoHugging Face20cua-lite /Lite.ScaleCUA cua-lite/Lite.ScaleCUA Lite.ScaleCUA grounded teacher trajectories collected on ScaleCUA's OSWorld tasks and judges via the cua-lite lite.scalecua runtime, from two teachers published as separate configs (*.gpt5_5 from gpt-5.5, *.qwen3_8_27b from Qwen/Qwen3.8-27B) and annotated by the same quality pass; ordinary quality gates tagged in metadata.others.exclude_reason, publish-invalid tool leaks/OOB coordinates hard-dropped (filter with not exclude_reason and episode_return>0.5)… See the full description on the dataset page: https://huggingface.co/datasets/cua-lite/Lite.ScaleCUA.imageimage-text-to-text100K<n<1M1 likes2.8k downloads10d agoHugging Face211aurent /unsplash-lite The Unsplash Lite Dataset (v1.2.1) The Lite dataset contains all of the same fields as the Full dataset, but is limited to ~25,000 photos. It can be used for both commercial and non-commercial usage, provided you abide by the terms. The Unsplash Dataset is made available for research purposes. It cannot be used to redistribute the images contained within. To use the Unsplash library in a product, see the Unsplash API. texttext-to-image10K<n<100K9 likes2.7k downloads3y agoHugging Face22sam-paech /livecodebench-code_generation_litetext1K<n<10K0 likes2.7k downloads1y agoHugging Face23CohereLabs /include-lite-44 INCLUDE-lite (44 languages) Dataset Description Paper: http://arxiv.org/abs/2411.19799 Dataset Summary INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed. It contains 11,095 4-option multiple-choice-questions (MCQ) extracted from academic and professional exams, covering 57 topics, including regional… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/include-lite-44.textmultiple-choice10K<n<100K16 likes2.5k downloads1y agoHugging Face24initiacms /XLRS-Bench-lite_VLM 🐙GitHub Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench 📜Dataset License Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from: DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply). ITCVDLicensed under CC-BY-NC-SA-4.0. MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench-lite_VLM.textvisual-question-answering1K<n<10K0 likes2.5k downloads11mo agoHugging Face25GSMA /ot-lite Open Telco Sample Data 1,850 telecom-specific evaluation samples across 8 benchmarks — designed for fast iteration during model development. Use this dataset for fast iteration during model development. Evaluate against GSMA/ot-full for final results. Eval Framework | Full Benchmarks Benchmarks | Config | Samples | Task | Paper | |--------|--------:|------|-------| | teleqna | 1,000 | Multiple-choice Q&A on telecom standards | arXiv | | teletables | 100 | Table… See the full description on the dataset page: https://huggingface.co/datasets/GSMA/ot-lite.textquestion-answering1K<n<10K2 likes2.5k downloads6mo agoHugging Face26harimo /scorio-lite Scorio Lite contains 1,211,520 sampled attempts from four model configurations and six reasoning benchmarks. Each model was run 80 times on every question. The five competition-math splits contain 186 questions. The superGPQA split contains a frozen, field-balanced sample of 3,600 questions. Each row includes the generation, rule-based grading, scores from CompassVerifier-3B and a reference-free verifier, and aggregate token statistics. Per-model configs also include token strings, log… See the full description on the dataset page: https://huggingface.co/datasets/harimo/scorio-lite.tabulartext-generation1M<n<10M0 likes2.3k downloads26d agoHugging Face27ArtificialAnalysis /AA-Briefcase-Lite AA-Briefcase-Lite The public example scenario for AA-Briefcase, Artificial Analysis' frontier agentic evaluation of realistic, long-horizon knowledge work. Leaderboard and detailed results Launch article AA-Briefcase extends frontier model benchmarking beyond coding and short-form reasoning to the professional deliverables knowledge workers produce day to day. It consists of four private scenarios in which agents complete realistic professional workflows across data science… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-Briefcase-Lite.documentothern<1K11 likes2.1k downloads3mo agoHugging Face28cua-lite /Lite.CUAGym cua-lite/Lite.CUAGym Lite.CUAGym grounded teacher trajectories collected on CUA-Gym task bundles and reward functions via the cua-lite lite.cuagym runtime, from two teachers published as separate configs (*.gpt5_5 from gpt-5.5, *.qwen3_8_27b from Qwen/Qwen3.8-27B) and annotated by the same quality pass; trajectories kept except /opt/env and OOB-coordinate hard-drops, quality gates tagged in metadata.others.exclude_reason (filter with not exclude_reason and episode_return>0.5)… See the full description on the dataset page: https://huggingface.co/datasets/cua-lite/Lite.CUAGym.imageimage-text-to-text10K<n<100K0 likes2.1k downloads12d agoHugging Face29MogoAI /Magpie_lite Magpie Dataset Lite Paper: Magpie: Real-Time World Renderer for Interactive GamesProject Page: https://zhanxy.xyz/Magpie-website Magpie Dataset Lite is a publicly released subset of the Magpie interactive game rendering dataset (arXiv:2608.27168). Magpie is a real-time generative world-rendering system that separates gameplay execution in a game engine from visual synthesis in a render server. This lite release provides 561 gameplay trajectories with a combined render.mp4… See the full description on the dataset page: https://huggingface.co/datasets/MogoAI/Magpie_lite.videovideo-to-video1K<n<10K1 likes2k downloads19d agoHugging Face30image-search-2 /unsplash_lite_image_dataset The Unsplash Dataset The Unsplash Dataset is made up of over 250,000+ contributing global photographers and data sourced from hundreds of millions of searches across a nearly unlimited number of uses and contexts. Due to the breadth of intent and semantics contained within the Unsplash dataset, it enables new opportunities for research and learning. The Unsplash Dataset is offered in two datasets: the Lite dataset: available for commercial and noncommercial usage, containing 25k… See the full description on the dataset page: https://huggingface.co/datasets/image-search-2/unsplash_lite_image_dataset.3 likes1.5k downloads5y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.