datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
spotify-tracks-lite
Context
This dataset consists of 24000 tracks from 30 genres, and is a shrunk version of maharshipandya/spotify-tracks-dataset dataset. All non-heuristic data is cut and cleaned for better usability and performance.
All data taken from Spotify API and is open source.
This dataset can be used to train prediction models based on user preferences, or categorise tracks by corresponding heuristic.
Column Description
danceability: Danceability describes how suitable a track is… See the full description on the dataset page: https://huggingface.co/datasets/engels/spotify-tracks-lite.protenix-data
Protenix Data
Protenix is ByteDance's open-source PyTorch reproduction of AlphaFold3, a biomolecular structure predictor that handles proteins, DNA, RNA, ligands, ions, and modifications under a unified all-atom diffusion model. Alongside the model code and weights, the team released the full preprocessed training dataset used to train Protenix and its successors, making it one of the largest publicly available AF3-style training corpora.
The released data is built from the wwPDB… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/protenix-data.vn-provinces-literacy-rate-age-15-plus
Vietnam provinces literacy rate (age 15+)
Literacy rate of population aged 15+ (percent). Coverage 2006 and 2009-2024. Includes historical Ha Tay in 2006. Year 2024 is preliminary. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison
Color key
Files
provinces… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-literacy-rate-age-15-plus.human-telemetry-driving-dataset-lite-version
Dataset Card for 15 Laps of 30Hz NGSIM-Style Telemetry
This is a Lite Version of a larger research dataset focusing on human driving signatures in high-fidelity simulations. It includes 15 full laps of telemetry captured at 30Hz within Unreal Engine 5, specifically formatted to match NGSIM standards.
Dataset Details
Dataset Description
This Lite Version dataset contains 15 laps of high-fidelity human driving telemetry. It is intended for researchers and… See the full description on the dataset page: https://huggingface.co/datasets/AtlasBuiltIt/human-telemetry-driving-dataset-lite-version.OpenProteinSet
OpenProteinSet
OpenProteinSet is an open-source corpus released by the OpenFold team (Ahdritz et al., NeurIPS 2023 Datasets and Benchmarks) that reproduces and extends the kind of training data used for AlphaFold2, which DeepMind never released. It contains more than 16 million precomputed multiple sequence alignments (MSAs), structural template hits from the Protein Data Bank, and AlphaFold2 structure predictions, and was used to train OpenFold from scratch to parity with… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/OpenProteinSet.real-toxicity-prompts-liteThis is a fork of the original RealToxicityPrompts dataset that contains a much smaller subset of the 100k prompts.
Subsets:
50_pct: This subset contains all the challenging prompts + 50% of the full RealToxicityPrompts size sampled from the other prompts.
10_pct: This subset contains all the challenging prompts + 10% of the full RealToxicityPrompts size sampled from the other prompts.
Please refer to the original dataset for the Dataset Card.
git_good_bench-lite
Dataset Summary
GitGoodBench Lite is a subset of 120 samples for evaluating the performance of AI agents in resolving git tasks (see Supported Scenarios).
The samples in the dataset are evenly split across the programming languages Python, Java and Kotlin and the sample types merge conflict resolution and file-commit gram.
This dataset thus contains 20 samples per sample type and programming language.
All data in this dataset are collected from 100 unique, open-source GitHub… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains/git_good_bench-lite.scientific-literature-research-assistant-dataEvolutionary
Evolutionary MSA Data
This repository contains precomputed evolutionary sequence-alignment data in an archive format that is practical to host and download from the Hub. The original file paths are preserved inside the tar shard, while metadata.csv gives a searchable index of every file.
The dataset is meant for workflows that need ready-to-use MSA/cache files without rebuilding them from sequence databases.
Contents
Component
Files
Size
msa_cache/
134,898… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/Evolutionary.bouncerbench-lite
Dataset Summary
Existing LLM-based tools and coding agents respond to every issue and generate a patch for every case, even when the input is vague or their own output is incorrect. There are no mechanisms in place to abstain when confidence is low. BouncerBench checks if AI agents know when not to act.
This is one of 3 datasets released as part of the paper Is Your Automated Software Engineer Trustworthy?.
input_bouncerTasks on bug‐report text. The model decides if a report is… See the full description on the dataset page: https://huggingface.co/datasets/uw-swag/bouncerbench-lite.test_thai_literaturegemini-2.0-flash-lite-pneumonia-datasetpatriae-cuba-literature-dataset
Patriae Cuban Literature Dataset (31k)
Dataset de literatura cubana curado por el equipo de Patriae como parte de su participación en el evento SomosNLP 2026, con el objetivo de emplearse por el mismo en la realización de tareas de reproducción del dialecto cubano.
📌 Nota de procedencia: Este repositorio es un espejo (mirror) oficial para el evento. El desarrollo activo, las actualizaciones del dataset y la autoría principal pertenecen a… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2026/patriae-cuba-literature-dataset.patriae-cuban-literature-dataset
Patriae Cuban Literature Dataset
Dataset de literatura cubana con +31k registros en formato CSV y Parquet, diseñado para el entrenamiento y evaluación de Modelos de Lenguaje (LLMs) y sistemas de Síntesis de Voz (TTS) orientados al dialecto y la cultura cubana.
Autores
Carlos Luis Barnés Infante (https://huggingface.co/blacknoize404)
Yisel Clavel Quintero (https://huggingface.co/clavel)
Curado por: Carlos Luis Barnés Infante
Descripción
Patriae… See the full description on the dataset page: https://huggingface.co/datasets/Patriae/patriae-cuban-literature-dataset.literacy-reading
