datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Workspace-Bench-Lite
Workspace-Bench-Lite
A Lightweight Subset of Workspace-Bench for Fast and Cost-Efficient Evaluation
Overview •
LeaderBoard •
Distribution •
Quick Start •
Changelog •
Citation
Overview
Workspace-Bench-Lite is the Lite split of Workspace-Bench 1.0, designed for fast iteration and lower-cost benchmarking while preserving the core evaluation setting of the full benchmark.
It contains 100 tasks selected from the full Workspace-Bench and is intended to… See the full description on the dataset page: https://huggingface.co/datasets/Workspace-Bench/Workspace-Bench-Lite.Minecraft-winter-shaders-Lite
❄️ Minecraft Winter Shaders (Lite)
10,000+ winter Minecraft screenshots with Photon Shaders.
For generative models, style transfer, and high-quality vision tasks.
Structure
Single folder: winter/
Format: 960x540 JPEG
License
Dataset: CC BY NC 4.0.Shader: Photon Shaders by SixthSurge (commercial use of screenshots explicitly allowed).See PHOTON_SHADERS_LICENSE.txt.
spotify-tracks-lite
Context
This dataset consists of 24000 tracks from 30 genres, and is a shrunk version of maharshipandya/spotify-tracks-dataset dataset. All non-heuristic data is cut and cleaned for better usability and performance.
All data taken from Spotify API and is open source.
This dataset can be used to train prediction models based on user preferences, or categorise tracks by corresponding heuristic.
Column Description
danceability: Danceability describes how suitable a track is… See the full description on the dataset page: https://huggingface.co/datasets/engels/spotify-tracks-lite.JP-TH_Literary_Translation_URL_Alignment_Index
JP–TH Literary Translation URL Alignment Index
This release provides a copyright-conscious metadata index and reproducibility package for a Japanese–Thai literary translation dataset associated with the study Context-Aware Prompting for Japanese–Thai Literary Translation in a Low-Resource Setting.
Overview
The release is designed to support reproducible academic research on Japanese–Thai literary machine translation, context-aware prompting, prompt engineering… See the full description on the dataset page: https://huggingface.co/datasets/Gsk068/JP-TH_Literary_Translation_URL_Alignment_Index.protenix-data
Protenix Data
Protenix is ByteDance's open-source PyTorch reproduction of AlphaFold3, a biomolecular structure predictor that handles proteins, DNA, RNA, ligands, ions, and modifications under a unified all-atom diffusion model. Alongside the model code and weights, the team released the full preprocessed training dataset used to train Protenix and its successors, making it one of the largest publicly available AF3-style training corpora.
The released data is built from the wwPDB… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/protenix-data.SWE-bench_Liteearnings_call_transcript_litehuman-telemetry-driving-dataset-lite-version
Dataset Card for 15 Laps of 30Hz NGSIM-Style Telemetry
This is a Lite Version of a larger research dataset focusing on human driving signatures in high-fidelity simulations. It includes 15 full laps of telemetry captured at 30Hz within Unreal Engine 5, specifically formatted to match NGSIM standards.
Dataset Details
Dataset Description
This Lite Version dataset contains 15 laps of high-fidelity human driving telemetry. It is intended for researchers and… See the full description on the dataset page: https://huggingface.co/datasets/AtlasBuiltIt/human-telemetry-driving-dataset-lite-version.vn-provinces-literacy-rate-age-15-plus
Vietnam provinces literacy rate (age 15+)
Literacy rate of population aged 15+ (percent). Coverage 2006 and 2009-2024. Includes historical Ha Tay in 2006. Year 2024 is preliminary. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison
Color key
Files
provinces… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-literacy-rate-age-15-plus.real-toxicity-prompts-liteThis is a fork of the original RealToxicityPrompts dataset that contains a much smaller subset of the 100k prompts.
Subsets:
50_pct: This subset contains all the challenging prompts + 50% of the full RealToxicityPrompts size sampled from the other prompts.
10_pct: This subset contains all the challenging prompts + 10% of the full RealToxicityPrompts size sampled from the other prompts.
Please refer to the original dataset for the Dataset Card.
OpenProteinSet
OpenProteinSet
OpenProteinSet is an open-source corpus released by the OpenFold team (Ahdritz et al., NeurIPS 2023 Datasets and Benchmarks) that reproduces and extends the kind of training data used for AlphaFold2, which DeepMind never released. It contains more than 16 million precomputed multiple sequence alignments (MSAs), structural template hits from the Protein Data Bank, and AlphaFold2 structure predictions, and was used to train OpenFold from scratch to parity with… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/OpenProteinSet.git_good_bench-lite
Dataset Summary
GitGoodBench Lite is a subset of 120 samples for evaluating the performance of AI agents in resolving git tasks (see Supported Scenarios).
The samples in the dataset are evenly split across the programming languages Python, Java and Kotlin and the sample types merge conflict resolution and file-commit gram.
This dataset thus contains 20 samples per sample type and programming language.
All data in this dataset are collected from 100 unique, open-source GitHub… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains/git_good_bench-lite.Kazakh-Literature-Collectionscientific-literature-research-assistant-datamodified_erotic_literature_collectionEvolutionary
Evolutionary MSA Data
This repository contains precomputed evolutionary sequence-alignment data in an archive format that is practical to host and download from the Hub. The original file paths are preserved inside the tar shard, while metadata.csv gives a searchable index of every file.
The dataset is meant for workflows that need ready-to-use MSA/cache files without rebuilding them from sequence databases.
Contents
Component
Files
Size
msa_cache/
134,898… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/Evolutionary.HanScript2HanViet2ModernViet_Literature
Use this dataset
You can load this dataset directly in your Python code using the 🤗 datasets library:
from datasets import load_dataset
dataset = load_dataset("GooglyEyeSuperman/HanScript2HanViet2ModernViet_Literature", {config_name})
config_name should be one of [train_official, train_comment, train_ai, test_official, dictionary, full_raw], to get the corresponding csv files.
Sino-Viet to Modern Vietnamese Dataset
Overview
This dataset focuses on… See the full description on the dataset page: https://huggingface.co/datasets/GooglyEyeSuperman/HanScript2HanViet2ModernViet_Literature.humanoid-pose-state-dataset-lite
Humanoid Pose State Dataset Lite
Lightweight synthetic dataset for humanoid robot pose classification.
Pose Classes
neutral
walking_pose
running_pose
sitting_pose
lifting_pose
waving_pose
Structure
dataset/
├── train/
├── validation/
Each split contains pose-labeled image folders.
Total Samples
Train: 600
Validation: 150
Image Format
RGB, 224x224
License
MIT
global-mmlu-lite
Global MMLU Lite – Galician & Urdu
Machine-translated Galician and Urdu subsets of the Global MMLU Lite benchmark.
Dataset Description
Global MMLU Lite is a culturally-aware, multilingual evaluation benchmark for large language models, covering multiple-choice questions across many academic subjects. This repository contains Galician (gl) and Urdu (ur) translations. This dataset was translated using Google Machine Translate.
Splits
Config… See the full description on the dataset page: https://huggingface.co/datasets/Owos/global-mmlu-lite.bouncerbench-lite
Dataset Summary
Existing LLM-based tools and coding agents respond to every issue and generate a patch for every case, even when the input is vague or their own output is incorrect. There are no mechanisms in place to abstain when confidence is low. BouncerBench checks if AI agents know when not to act.
This is one of 3 datasets released as part of the paper Is Your Automated Software Engineer Trustworthy?.
input_bouncerTasks on bug‐report text. The model decides if a report is… See the full description on the dataset page: https://huggingface.co/datasets/uw-swag/bouncerbench-lite.statistical_literacyv2test_thai_literaturepatriae-cuban-literature-dataset
Patriae Cuban Literature Dataset (40k)
Dataset de literatura cubana con 40k registros en formato CSV y Parquet, diseñado para el entrenamiento y evaluación de Modelos de Lenguaje (LLMs) y sistemas de Síntesis de Voz (TTS) orientados al dialecto y la cultura cubana.
Autores
Carlos Luis Barnés Infante (https://huggingface.co/blacknoize404)
Yisel Clavel Quintero (https://huggingface.co/clavel)
Curado por: Carlos Luis Barnés Infante
Descripción del… See the full description on the dataset page: https://huggingface.co/datasets/blacknoize404/patriae-cuban-literature-dataset.hdfs-log-anomaly-dataset-lite
Lightweight Transformer for HDFS Log Anomaly Detection
This repository provides the processed benchmark dataset and preprocessing scripts
used in our study on HDFS log anomaly detection with lightweight Transformer models.
Contents
data/: Benchmark-ready train and test datasets
preprocessing/: Scripts for generating the processed datasets
Reproducibility
The released datasets correspond exactly to the experimental setup described in the
paper. Model training… See the full description on the dataset page: https://huggingface.co/datasets/ridwanakm/hdfs-log-anomaly-dataset-lite.dataset-20251211-lite
dataset-20251211-lite
Created on: 2025-12-11T13:36:13.447423+00:00
Session ID: 2025-12-11T13:36:13.447423+00:00-2357
literaturemmtg_liteliterary-reasoning-filteredFiltered version of the [https://huggingface.co/datasets/agentlans/literary-reasoning] dataset with only english entries where genre is not NULL.
dataset-20251216-lite
dataset-20251216-lite
Created on: 2025-12-16T04:53:12.979733+00:00
Session ID: 2025-12-16T04:53:12.979733+00:00-7911
ada_igbo_literacy_dataset
