datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DecodingTrust-Agent-Platform
DecodingTrust-Agent Platform
A Controllable and Interactive Red-Teaming Platform for AI Agents.
This is the per-task dataset for the DecodingTrust-Agent Platform (DTAP),
spanning 14 real-world domains and 50+ simulation environments that replicate widely-used
systems such as Google Workspace, PayPal, Slack, Salesforce, Snowflake, and Databricks. Each task
ships the configuration the evaluator needs to spin up the sandbox, run an agent, and verify the
outcome — config.yaml (task… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/DecodingTrust-Agent-Platform.brain-decodingDecodingTrust
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Overview
This repo contains the source code of DecodingTrust. This research endeavor is designed to help researchers better understand the capabilities, limitations, and potential risks associated with deploying these state-of-the-art Large Language Models (LLMs). See our paper for details.
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Boxin Wang, Weixin Chen, Hengzhi… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/DecodingTrust.DLM-Decoding-Analysis
DLM-Decoding-Analysis
Diffusion Language Model Knows the Answer Before It Decodes
Pengxiang Li*, Yefan Zhou*, Dilxat Muhtar, Lu Yin, Shilin Yan, Li Shen, Yi Liang, Soroush Vosoughi, Shiwei Liu
The Fourteenth International Conference on Learning Representations (ICLR 2026)
TL;DR: Diffusion language models often commit to the correct answer
well before they finish decoding. This dataset releases the per-question,
step-by-step decoding trajectories of LLaDA-8B-Instruct on… See the full description on the dataset page: https://huggingface.co/datasets/YefanZhou98/DLM-Decoding-Analysis.speculative_decoding_benchmarksbrain-decoding-nsddecodingtrust-windows-qcow2ca1-position-decoding
ca1-position-decoding
Data for the terminal-bench-science task ca1-position-decoding: decode a mouse's position in an
open field from raw two-photon calcium imaging of hippocampal CA1. This card is the only place the
provenance is written down; the task deliberately gives the agent no acquisition metadata beyond
the frame rate, the pixel scales and the plane alternation stated in its instruction.
Source
Zong, W., Obenhaus, H. A., Skytoen, E. R., et al. (2022).… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/ca1-position-decoding.decodingthoughtsdecoding-robustness-results
Decoding Robustness Results
Mechanistic robustness evaluation results for language models under six input
perturbations: character replacement, BPE-token replacement, word replacement,
local token shuffle, typographical corruption, and synonym replacement.
The repository is organized by model and perturbation:
models/<model>/<perturbation>/<percentage>/evals.csv
The qwen2.5_1.5b/adversarial directory contains the separate adversarial
evaluation outputs and manifest. Failed or… See the full description on the dataset page: https://huggingface.co/datasets/christian-hoang-04/decoding-robustness-results.speculative-decoding-bench-rtx4090
Speculative Decoding Benchmark — RTX 4090
TL;DR: 4,576 benchmark runs measuring speculative decoding speedup / acceptance rate
across llama.cpp and LM Studio, Qwen3 (8B/14B) and Llama-3.1-8B target models, on a
single consumer RTX 4090 (24GB). Best observed case: the draft-free ngram-mod
self-speculative mode on structured tasks (JSON extraction 2.81x, code 2.76x,
global-median aggregation at temp=0). Open-ended tasks (creative writing, translation)
with a traditional draft… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/speculative-decoding-bench-rtx4090.speculative-decoding-benchmark-resultsdecodingtrust-macos-qcow2eagle3-speculative-decoding-energy-sweep
EAGLE3 Speculative Decoding Energy Sweep
Per-config energy/throughput/latency measurements for EAGLE3 speculative decoding
(speculative_num_steps, speculative_eagle_topk, speculative_num_draft_tokens)
served with sglang, across batch sizes. Collected for an RL project that learns to
pick speculative-decoding parameters to hold GPU energy utilization in a target band.
Model: unsloth/Llama-3.2-1B-Instruct + rescommons/SpecForge-EAGLE3-Llama-3.2-1B-Instruct draft head.
Hardware:… See the full description on the dataset page: https://huggingface.co/datasets/Pradheep1647/eagle3-speculative-decoding-energy-sweep.speculative-decoding-papers
Speculative Decoding Papers — FineSet
A research-paper dataset on Speculative Decoding Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on Speculative Decoding Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset
Quality-scored:… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/speculative-decoding-papers.ml-lecture-2021-longDerived from: ky552/ML2021_ASR_ST
Segments from the same lecture are concatenated together.
repro-entropy-informed-decoding-adaptive-information-driven-branching-traces
Agent traces
Agent sessions published from a Trackio Logbook.
decodingtrust-windows-filesai4sci-surface-code-decoding
Sycamore surface-code decoding: materialized benchmark
Predict a logical observable flip from repeated stabilizer detection events in
a noisy quantum memory. The benchmark trains decoders that improve the
reliability of encoded quantum information. It uses real Sycamore hard-readout
experiments at code distances 3 and 5, not simulated soft-readout d11 data.
Source: Google Quantum AI Sycamore memory experiments, Zenodo 6804040,
CC-BY-4.0. Scientific model reference:
Bausch et… See the full description on the dataset page: https://huggingface.co/datasets/Corning/ai4sci-surface-code-decoding.clean_squad_v1
Clean SQuAD v1
This is a refined version of the SQuAD v1 dataset. It has been preprocessed to ensure higher data quality and usability for NLP tasks such as Question Answering.
Description
The Clean SQuAD v1 dataset was created by applying preprocessing steps to the original SQuAD v1 dataset, including:
Trimming whitespace: All leading and trailing spaces have been removed from the question field.
Minimum question length: Questions with fewer than 12 characters were… See the full description on the dataset page: https://huggingface.co/datasets/decodingchris/clean_squad_v1.message-decoding-words-and-sequences-r1decoding_llama3HH_gemma-2-2b-it
Helpful-Harmless Dataset with Responses Generated from gemma-2-2b-it
This dataset is used to train the value functions and test methods in Robust Multi-Objective Decoding paper.
We take the prompts taken from Helpful-Harmless dataset (Bai et al., 2022), and use gemma-2-2b-it to generate 4 responses per prompt. Each response is generated up to 256 tokens.
Each response is evaluated with Ray2333/gpt2-large-helpful-reward_model and Ray2333/gpt2-large-harmless-reward_model.
ai4sci-surface-code-decoding
Sycamore surface-code decoding: materialized benchmark
Predict a logical observable flip from repeated stabilizer detection events in
a noisy quantum memory. The benchmark trains decoders that improve the
reliability of encoded quantum information. It uses real Sycamore hard-readout
experiments at code distances 3 and 5, not simulated soft-readout d11 data.
Source: Google Quantum AI Sycamore memory experiments, Zenodo 6804040,
CC-BY-4.0. Scientific model reference:
Bausch et… See the full description on the dataset page: https://huggingface.co/datasets/sunweiwei/ai4sci-surface-code-decoding.daily-paper-2026-09-25-spec-decoding-acceptance-output-structure
The Draft Law: Measuring How Agentic Output Structure Sets Speculative-Decoding Acceptance and Net Per-Token Cost on Self-Hosted H200
TL;DR — An analytical paper deriving the draft law for speculative decoding on agentic traffic over a self-hosted single H200: expected accepted prefix length is set by the mean structural predictability of the output - schema-constrained tool-call, reasoning, and prose positions - and the net per-token saving is strictly increasing in the… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-09-25-spec-decoding-acceptance-output-structure.ca1-online-decoding
ca1-online-decoding
Data for the terminal-bench-science task ca1-online-decoding: decode a mouse's position in an
open field from a stream of raw two-photon calcium imaging of hippocampal CA1, one frame at a time.
This card is the only place the provenance is written down; the task deliberately gives the agent
no acquisition metadata beyond the frame rate, the pixel scales and the plane alternation stated in
its instruction.
Source
Zong, W., Obenhaus, H. A.… See the full description on the dataset page: https://huggingface.co/datasets/RyanIRL/ca1-online-decoding.speculative-decoding-datasetmessage-decoding-abc-zoom-indecoding_summaries_temperature_0.4message-decoding-dataset
