datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_generation_liteLiveCodeBench is a temporaly updating benchmark for code generation. Please check the homepage: https://livecodebench.github.io/.SWE-bench_Lite
Dataset Summary
SWE-bench Lite is subset of SWE-bench, a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution.
The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Want to run inference now?
This dataset only contains the… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_Lite.R2E-Gym-LiteSWE-bench_Lite
Dataset Summary
SWE-bench Lite is subset of SWE-bench, a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution.
The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Want to run inference now?
This dataset only contains the… See the full description on the dataset page: https://huggingface.co/datasets/SWE-bench/SWE-bench_Lite.Global-MMLU-Lite
Releases:
Version 3.0 (May 2026): GMMLU Lite 3.0 release with 5 new languages: Czech, Hungarian, Italian (updated), Oriya, Slovak and Tajik
Version 2.0 (Dec 2025): GMMLU Lite 2.0 release with 3 new languages: Albanian, Burmese and Welsh
Version 1.0 (Dec 2024): GMMLU Lite initial release with 15 languages.
Dataset Summary
Global-MMLU-Lite is a multilingual evaluation set spanning 23 languages, including English. It is "lite" version of the original Global-MMLU… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/Global-MMLU-Lite.SWE-Gym-LiteSWE-Gym Lite contains 230 instances sourced from 11 Python repos, following SWE-Bench Lite data collection procedure.
Get started at project page github.com/SWE-Gym/SWE-Gym
LMMs-Eval-LiteXLRS-Bench-lite
🐙GitHub
Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench
📜Dataset License
Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from:
DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply).
ITCVDLicensed under CC-BY-NC-SA-4.0.
MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench-lite.MMLU-ProX-Lite
MMLU-ProX-Lite
MMLU-ProX is a multilingual benchmark that builds upon MMLU-Pro, extending to 29 typologically diverse languages, designed to evaluate large language models' reasoning capabilities across linguistic and cultural boundaries.
Github | Paper
News
[2025/08] 🎉 MMLU-ProX was accepted by EMNLP 2025 Main Conference!
[2025/05] MMLU-ProX now contains 29 languages, all available on Huggingface.
[2025/03] MMLU-ProX is now available on Huggingface.
[2025/03] We are… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/MMLU-ProX-Lite.SWE-bench_Lite_filteredErotic_Literature_CollectionEnglish
中文色情文学数据集合集
概述
本仓库包含了51个中文色情文学数据集。每个数据集由短篇色情小说、个人色情经验及其他形式的色情内容组成。数据集的格式为JSON,每个文件包含一个对象数组,每个对象代表一篇文档:
[
{"text": "document"},
{"text": "document"}
]
这些数据集可用于语言模型的预训练,经过适当调整后也可用于模型的微调。
数据集格式
文件格式: JSON
内容: 短篇色情小说、个人色情经验及其他色情内容
结构:
每个文件包含一个对象数组
每个对象包含一个键 "text",其值为相应的文档内容
使用方法
这些数据集主要用于研究目的,特别是在语言模型的开发和微调中使用。由于内容的敏感性,用户应谨慎处理这些数据集,并确保遵守当地的法律法规及相关指导原则。
示例用法
import json
# 加载数据集with open('path_to_json_file.json', 'r'… See the full description on the dataset page: https://huggingface.co/datasets/ystemsrx/Erotic_Literature_Collection.lite.cuaworld-assets
lite.cuaworld materials
Maintained environment materials for cua-lite's lite.cuaworld.* software
environments — forked from cmu-l3/gym-anything
(CUA-World, MIT) and curated/edited/expanded by us. Published as the Hugging Face dataset
cua-lite/lite.cuaworld-assets.
This repo holds content only (no engine code): per-environment env.json,
install/setup scripts/, per-task assets (tasks/<task>/…), data/, config/,
assets/, optional post_build.sh, and a curated registered.json. The… See the full description on the dataset page: https://huggingface.co/datasets/cua-lite/lite.cuaworld-assets.PDB
PDB mmCIF Entry Index
The Protein Data Bank is the single global archive of experimentally-determined 3D structures of biological macromolecules, established in 1971 and now holding well over 230,000 entries. It stores atomic coordinates for proteins, nucleic acids, and their complexes determined by X-ray crystallography, cryo-EM, NMR, micro-electron diffraction, and integrative methods, along with the underlying experimental data (structure factors, EM maps, NMR restraints) and… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/PDB.code_generation_lite
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
🏠 Home Page •
💻 GitHub Repository •
🏆 Leaderboard •
📄 Paper
Change Log
Since LiveCodeBench is a continuously updated benchmark, we provide different versions of the dataset. Particularly, we provide the following versions of the dataset:
release_v1: The initial release of the dataset with problems released between May 2023 and Mar 2024 containing 400… See the full description on the dataset page: https://huggingface.co/datasets/lighteval/code_generation_lite.Workspace-Bench-Lite
Workspace-Bench-Lite
A Lightweight Subset of Workspace-Bench for Fast and Cost-Efficient Evaluation
Overview •
LeaderBoard •
Distribution •
Quick Start •
Changelog •
Citation
Overview
Workspace-Bench-Lite is the Lite split of Workspace-Bench 1.0, designed for fast iteration and lower-cost benchmarking while preserving the core evaluation setting of the full benchmark.
It contains 100 tasks selected from the full Workspace-Bench and is intended to… See the full description on the dataset page: https://huggingface.co/datasets/Workspace-Bench/Workspace-Bench-Lite.SWE-bench_Lite_filtered_1book2-lite-cleanedLiterature-zhA composite of Chinese books, papers, legal documents, and patents from common crawl.
Data cleaning
Text with more than 2% of non-Latin, non-Chinese characters are removed.
Text with large portions of special characters are removed.
Traditional Chinese is converted to simplified Chinese.
Model filtering
Qwen2.5-32B-Instruct is used to generate language quality annotation (on a scale of 1-5) for 398K Chinese samples and 250K English samples. An XLM-RoBERT-large classifier is trained with… See the full description on the dataset page: https://huggingface.co/datasets/Geralt-Targaryen/Literature-zh.financial-analyst-data-lite
financial-analyst-data-lite
EN: A-share historical OHLCV + valuation + financials + TDX F10 events, packaged in Qlib binary + Parquet formats. Companion dataset for financial-analyst — a 14-agent single-stock deep-dive research workstation.
中文: A 股历史行情 + 估值 + 财报 + TDX F10 事件数据集, Qlib 二进制 + Parquet 双格式打包. 配套 financial-analyst — 14 Agent 个股深度研究工作站使用.
Published / 发布: 2026-05-24 · Size / 体量: ~2.72 GB · License: Apache 2.0
📊 Three Preset Tiers / 三档预设
Pick the tier that… See the full description on the dataset page: https://huggingface.co/datasets/yifishbossman/financial-analyst-data-lite.Lite.ScaleCUA
cua-lite/Lite.ScaleCUA
Lite.ScaleCUA grounded teacher trajectories collected on ScaleCUA's OSWorld tasks and judges via the cua-lite lite.scalecua runtime, from two teachers published as separate configs (*.gpt5_5 from gpt-5.5, *.qwen3_8_27b from Qwen/Qwen3.8-27B) and annotated by the same quality pass; ordinary quality gates tagged in metadata.others.exclude_reason, publish-invalid tool leaks/OOB coordinates hard-dropped (filter with not exclude_reason and episode_return>0.5)… See the full description on the dataset page: https://huggingface.co/datasets/cua-lite/Lite.ScaleCUA.unsplash-lite
The Unsplash Lite Dataset (v1.2.1)
The Lite dataset contains all of the same fields as the Full dataset, but is limited to ~25,000 photos.
It can be used for both commercial and non-commercial usage, provided you abide by the terms.
The Unsplash Dataset is made available for research purposes.
It cannot be used to redistribute the images contained within.
To use the Unsplash library in a product, see the Unsplash API.
livecodebench-code_generation_liteinclude-lite-44
INCLUDE-lite (44 languages)
Dataset Description
Paper: http://arxiv.org/abs/2411.19799
Dataset Summary
INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed.
It contains 11,095 4-option multiple-choice-questions (MCQ) extracted from academic and professional exams, covering 57 topics, including regional… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/include-lite-44.XLRS-Bench-lite_VLM
🐙GitHub
Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench
📜Dataset License
Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from:
DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply).
ITCVDLicensed under CC-BY-NC-SA-4.0.
MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench-lite_VLM.ot-lite
Open Telco Sample Data
1,850 telecom-specific evaluation samples across 8 benchmarks — designed for fast iteration during model development.
Use this dataset for fast iteration during model development. Evaluate against GSMA/ot-full for final results.
Eval Framework | Full Benchmarks
Benchmarks
| Config | Samples | Task | Paper |
|--------|--------:|------|-------|
| teleqna | 1,000 | Multiple-choice Q&A on telecom standards | arXiv |
| teletables | 100 | Table… See the full description on the dataset page: https://huggingface.co/datasets/GSMA/ot-lite.scorio-lite
Scorio Lite contains 1,211,520 sampled attempts from four model configurations and
six reasoning benchmarks. Each model was run 80 times on every question.
The five competition-math splits contain 186 questions. The superGPQA split contains a
frozen, field-balanced sample of 3,600 questions. Each row includes the generation,
rule-based grading, scores from CompassVerifier-3B and a reference-free verifier, and
aggregate token statistics. Per-model configs also include token strings, log… See the full description on the dataset page: https://huggingface.co/datasets/harimo/scorio-lite.AA-Briefcase-Lite
AA-Briefcase-Lite
The public example scenario for AA-Briefcase, Artificial Analysis' frontier agentic evaluation of realistic, long-horizon knowledge work.
Leaderboard and detailed results
Launch article
AA-Briefcase extends frontier model benchmarking beyond coding and short-form reasoning to the professional deliverables knowledge workers produce day to day. It consists of four private scenarios in which agents complete realistic professional workflows across data science… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-Briefcase-Lite.Lite.CUAGym
cua-lite/Lite.CUAGym
Lite.CUAGym grounded teacher trajectories collected on CUA-Gym task bundles and reward functions via the cua-lite lite.cuagym runtime, from two teachers published as separate configs (*.gpt5_5 from gpt-5.5, *.qwen3_8_27b from Qwen/Qwen3.8-27B) and annotated by the same quality pass; trajectories kept except /opt/env and OOB-coordinate hard-drops, quality gates tagged in metadata.others.exclude_reason (filter with not exclude_reason and episode_return>0.5)… See the full description on the dataset page: https://huggingface.co/datasets/cua-lite/Lite.CUAGym.Magpie_lite
Magpie Dataset Lite
Paper: Magpie: Real-Time World Renderer for Interactive GamesProject Page: https://zhanxy.xyz/Magpie-website
Magpie Dataset Lite is a publicly released subset of the Magpie interactive game rendering dataset (arXiv:2608.27168). Magpie is a real-time generative world-rendering system that separates gameplay execution in a game engine from visual synthesis in a render server. This lite release provides 561 gameplay trajectories with a combined render.mp4… See the full description on the dataset page: https://huggingface.co/datasets/MogoAI/Magpie_lite.unsplash_lite_image_dataset
The Unsplash Dataset
The Unsplash Dataset is made up of over 250,000+ contributing global photographers and data sourced from hundreds of millions of searches across a nearly unlimited number of uses and contexts. Due to the breadth of intent and semantics contained within the Unsplash dataset, it enables new opportunities for research and learning.
The Unsplash Dataset is offered in two datasets:
the Lite dataset: available for commercial and noncommercial usage, containing 25k… See the full description on the dataset page: https://huggingface.co/datasets/image-search-2/unsplash_lite_image_dataset.
