datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
measuring-hate-speech
Dataset card for Measuring Hate Speech
This is a public release of the dataset described in Kennedy et al. (2020) and Sachdeva et al. (2022), consisting of 39,565 comments annotated by 7,912 annotators, for 135,556 combined rows. The primary outcome variable is the "hate speech score" but the 10 constituent ordinal labels (sentiment, (dis)respect, insult, humiliation, inferior status, violence, dehumanization, genocide, attack/defense, hate speech benchmark) can also be treated as… See the full description on the dataset page: https://huggingface.co/datasets/ucberkeley-dlab/measuring-hate-speech.dlgenai-nppe-datasetcorpus-1T-manifest
SPP Corpus 1T Manifest
The selection manifest for the ~1.0T-token pretraining corpus used in
Synthetic Persona Pretraining (SPP): Alignment from Token Zero.
The corpus is a seeded subsample of allenai/dolma3_mix-6T.
Rather than redistribute ~2.6 TB of text that is already public, this dataset
publishes the selection decisions keyed by upstream document id, so the corpus
can be reconstructed exactly by replaying against upstream.
📄 Reflections + text for the annotated half:… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/corpus-1T-manifest.reflection-50m
SPP Reflection 50M
The 51.4M-document reflection set from Synthetic Persona Pretraining (SPP):
Alignment from Token Zero — the production half-corpus run, and the dataset the
released models were actually trained on.
🔬 Small sample (same format): dlab-spp/reflection-sample-2k
📉 Earlier 10M run: dlab-spp/reflection-10m
🧾 Safety scores for the full 1T corpus: dlab-spp/safety-classifications
Each row pairs a source document with two generated constitution reflections — a… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-50m.dlgenai-nppe2-datasetdlam-ts-project-data-2026
operations_forecasting_2026
Multivariate hourly forecasting for anonymized operations units.
Target
Predict the future hourly operational load index for each series_id. Higher values indicate more operational pressure in that unit.
Forecast Contract
Frequency: h
Series: 96
Timesteps per series: 4992
Target column: target
Training history length used by the baseline templates: 168
Rollout block length: 24
Required prediction horizon: validation: 336, test: 336… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/dlam-ts-project-data-2026.DLD_Transactionsdlr_edan_shared_control_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "dlr_edan",
"total_episodes": 104,
"total_frames": 8928,
"total_tasks": 10,
"total_videos": 104,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:104"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/dlr_edan_shared_control_lerobot.dlr_edan_shared_controlThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 104,
"total_frames": 8928,
"total_tasks": 14,
"total_videos": 104,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:104"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/dlr_edan_shared_control.dlr_sara_grid_clampThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 107,
"total_frames": 7622,
"total_tasks": 1,
"total_videos": 107,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:107"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/dlr_sara_grid_clamp.reflection-10m
SPP Reflection 10M
The full ~10M-document reflection set from Synthetic Persona Pretraining (SPP):
Alignment from Token Zero.
📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero
🔬 Small sample (same format): dlab-spp/reflection-sample-2k — a 2,000-row sample drawn from this set, for quick inspection.
Each row pairs a pretraining document with a synthetic, value-laden reflection
generated for it: a short first-person (and third-person) moral reflection… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-10m.dlr_sara_pourThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 100,
"total_frames": 12971,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/dlr_sara_pour.tcga-dlbc-tabular-open
TCGA-DLBC — Tabular (Open Access)
Open-access TCGA-DLBC data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 03:53:50 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-dlbc-tabular-open.proj-dllm-sftplue
PLUE
Repository: https://github.com/ju-resplande/PLUE
Paper:
Leaderboard:
Point of Contact:
Portuguese translation of the GLUE benchmark, SNLI, and Scitail using OPUS-MT model and Google Cloud Translation.
The language data in PLUE is Brazilian Portuguese (BCP-47 pt-BR)
Citation Information
@misc{Gomes2020,
author = {GOMES, J. R. S.},
title = {PLUE: Portuguese Language Understanding Evaluation},
year = {2020},
publisher = {GitHub},
journal = {GitHub… See the full description on the dataset page: https://huggingface.co/datasets/dlb/plue.interaction_protocol
Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate
This repository contains the processed experimental datasets used in:
Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate. Pratik S. Sachdeva and Tom van Nuenen. COLM 2026.
Dataset contents
The experiments/ directory contains Parquet datasets used to reproduce the
figures and analyses in the paper. It includes:
synchronous head-to-head debates;
round-robin head-to-head debates;… See the full description on the dataset page: https://huggingface.co/datasets/ucberkeley-dlab/interaction_protocol.dllm-img-edit-vq-cachedual_so101This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "dual_so101_follower",
"total_episodes": 2,
"total_frames": 582,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dleon23/dual_so101.DLT-Tweets
DLT-Tweets
[Paper] •
[Code]
Dataset Description
Dataset Summary
DLT-Tweets is a large-scale corpus of social media posts related to Distributed Ledger Technology (DLT). This dataset is part of the larger DLT-Corpus collection, designed to support NLP research, social computing studies, and public discourse analysis in the DLT domain. It was introduced in the paper DLT-Corpus: A Large-Scale Text Collection for the Distributed Ledger Technology Domain.… See the full description on the dataset page: https://huggingface.co/datasets/ExponentialScience/DLT-Tweets.meteora-dlmm-historical-data
Meteora DLMM Historical Data
Decoded Solana mainnet instructions and events from Meteora DLMM (Dynamic Liquidity Market Maker), a concentrated-liquidity DEX where liquidity sits in discrete price bins and the fee rate rises with volatility.
74 tables, 59,575 rows, one row per decoded instruction or event. Program ID LBUZKhRxPF3XUpBCjp4YzTKgLccjZhTSDM9YuVaPwxo.
This is a free sample from datastore.sh, which publishes the complete history as versioned Parquet.
Read… See the full description on the dataset page: https://huggingface.co/datasets/DataStore/meteora-dlmm-historical-data.dl-trm-phase2-codebooks
DL-TRM Phase 2 Codebooks
This dataset repository contains Phase 2 transition VQ codebook artifacts for DL-TRM.
Contents are organized by vocabulary size:
V16/
V32/
V128/
V256/
Each folder includes:
z_traces.pt: discrete Z trace dataset for the Phase 1 route traces
codebook.pt: learned VQ codebook weights
transition_vq_model.pt: trained transition VQ model weights
diagnostics.json: code usage and final training diagnostics
z_trace_manifest.json: artifact manifest
checkpoints/:… See the full description on the dataset page: https://huggingface.co/datasets/omrisap/dl-trm-phase2-codebooks.dl-npppe2-datasetdl3dv_InP_480top_tagging
Dataset Card for Top Quark Tagging
Dataset Summary
Top Quark Tagging is a dataset of Monte Carlo simulated events produced by proton-proton collisions at the Large Hadron Collider. The top-quark signal and mixed quark-gluon background jets are produced with Pythia8 with its default tune for a center-of-mass energy of 14 TeV. Multiple interactions and pile-up are ignored. The leading 200 jet constituent four-momenta (E,px,py,pz) (E, p_x, p_y, p_z) (E,px,py,pz)are stored… See the full description on the dataset page: https://huggingface.co/datasets/dl4phys/top_tagging.wiki-sim
Wiki Sim
Overview
This new semi-synthetic dataset is derived from wikimedia/wikipedia.
Each row contains 1-3 references sentences extracted from the original dataset.
For each reference sentence, we use an optimized DSPy program to generate 4 similar sentences:
Synonym (Replace words with synonyms to maintain the same meaning.)
Paraphrase (Rephrase the sentence using a different structure while keeping the same idea.)
Conceptual Overlap (Express a related concept… See the full description on the dataset page: https://huggingface.co/datasets/dleemiller/wiki-sim.read_test_202608110744310438FineCat-NLI
Fine Concatenation (FineCat) NLI
Overview
A common criticism of SNLI and MNLI datasets is that there are too many 'easy' samples.
This tends to overfit to simple / trivial patterns that don't generalize well.
In order to combat this, I concatenated 7 datasets (2.6M samples), then ran a
training test for 50k steps with ModernBERT-large in cross-encoder configuration.
I found that 1 dataset (~100k samples) did not have good compatibility with the labels of the others… See the full description on the dataset page: https://huggingface.co/datasets/dleemiller/FineCat-NLI.UM-DLP-Public-Benchmarking-Dataset
UM DLP Public Benchmarking Dataset
Description
The UM DLP Public Benchmarking Dataset is a publicly available collection designed specifically to stress test Data Loss Prevention (DLP) systems, helping identify detection gaps, false positives, and false negatives for ongoing improvement.
This benchmark dataset contains 1,343 manually validated records across six major categories relevant to financial and sensitive data risks:
Financial Data (Account information about… See the full description on the dataset page: https://huggingface.co/datasets/alibustami/UM-DLP-Public-Benchmarking-Dataset.trl-dlte
TRL-DLTE
Paper: arXiv:2606.09323 — TRL-Bench: Standardizing Cross-Paradigm Representation-Level Evaluation of Tabular Encoders · Code: LOGO-CUHKSZ/TRL-Bench
Compositional Data-Lake Table Enrichment suite of TRL-Bench. A 47,772-table data lake derived from 1,379 TabFact and WikiTableQuestions parent tables, fragmented at four cumulative noise tiers (clean / schema / cell / hard). Each parent yields a seed query, a union target (additional rows), and a join target (additional… See the full description on the dataset page: https://huggingface.co/datasets/logo-lab/trl-dlte.pick_and_placeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "omx_follower",
"total_episodes": 40,
"total_frames": 25782,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dl1wjd2/pick_and_place.
