datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DecodingTrust
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Overview
This repo contains the source code of DecodingTrust. This research endeavor is designed to help researchers better understand the capabilities, limitations, and potential risks associated with deploying these state-of-the-art Large Language Models (LLMs). See our paper for details.
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Boxin Wang, Weixin Chen, Hengzhi… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/DecodingTrust.brain-decoding-nsdca1-position-decoding
ca1-position-decoding
Data for the terminal-bench-science task ca1-position-decoding: decode a mouse's position in an
open field from raw two-photon calcium imaging of hippocampal CA1. This card is the only place the
provenance is written down; the task deliberately gives the agent no acquisition metadata beyond
the frame rate, the pixel scales and the plane alternation stated in its instruction.
Source
Zong, W., Obenhaus, H. A., Skytoen, E. R., et al. (2022).… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/ca1-position-decoding.decodingthoughtsspeculative-decoding-bench-rtx4090
Speculative Decoding Benchmark — RTX 4090
TL;DR: 4,576 benchmark runs measuring speculative decoding speedup / acceptance rate
across llama.cpp and LM Studio, Qwen3 (8B/14B) and Llama-3.1-8B target models, on a
single consumer RTX 4090 (24GB). Best observed case: the draft-free ngram-mod
self-speculative mode on structured tasks (JSON extraction 2.81x, code 2.76x,
global-median aggregation at temp=0). Open-ended tasks (creative writing, translation)
with a traditional draft… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/speculative-decoding-bench-rtx4090.speculative-decoding-benchmark-resultseagle3-speculative-decoding-energy-sweep
EAGLE3 Speculative Decoding Energy Sweep
Per-config energy/throughput/latency measurements for EAGLE3 speculative decoding
(speculative_num_steps, speculative_eagle_topk, speculative_num_draft_tokens)
served with sglang, across batch sizes. Collected for an RL project that learns to
pick speculative-decoding parameters to hold GPU energy utilization in a target band.
Model: unsloth/Llama-3.2-1B-Instruct + rescommons/SpecForge-EAGLE3-Llama-3.2-1B-Instruct draft head.
Hardware:… See the full description on the dataset page: https://huggingface.co/datasets/Pradheep1647/eagle3-speculative-decoding-energy-sweep.speculative-decoding-papers
Speculative Decoding Papers — FineSet
A research-paper dataset on Speculative Decoding Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on Speculative Decoding Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset
Quality-scored:… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/speculative-decoding-papers.ml-lecture-2021-longDerived from: ky552/ML2021_ASR_ST
Segments from the same lecture are concatenated together.
clean_squad_v1
Clean SQuAD v1
This is a refined version of the SQuAD v1 dataset. It has been preprocessed to ensure higher data quality and usability for NLP tasks such as Question Answering.
Description
The Clean SQuAD v1 dataset was created by applying preprocessing steps to the original SQuAD v1 dataset, including:
Trimming whitespace: All leading and trailing spaces have been removed from the question field.
Minimum question length: Questions with fewer than 12 characters were… See the full description on the dataset page: https://huggingface.co/datasets/decodingchris/clean_squad_v1.message-decoding-words-and-sequences-r1decoding_llama3ca1-online-decoding
ca1-online-decoding
Data for the terminal-bench-science task ca1-online-decoding: decode a mouse's position in an
open field from a stream of raw two-photon calcium imaging of hippocampal CA1, one frame at a time.
This card is the only place the provenance is written down; the task deliberately gives the agent
no acquisition metadata beyond the frame rate, the pixel scales and the plane alternation stated in
its instruction.
Source
Zong, W., Obenhaus, H. A.… See the full description on the dataset page: https://huggingface.co/datasets/RyanIRL/ca1-online-decoding.speculative-decoding-datasetmessage-decoding-abc-zoom-indecoding_summaries_temperature_0.4message-decoding-datasetmessage-decoding-words-and-sequences-target-zoom-in-r1Decoding-Text-Summarization-Most-Frequent-Words-and-Medical-Text-Detectionmessage-decoding-words-and-sequences-target-zoom-indecoding_summaries_temperature_0.6decoding_summaries_temperature_0.7decoding_summaries_temperature_0.8decoding_summaries_top_k_50decoding_trustCreated by Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, Sang T. Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, Bo Li.
Bad-Decoding-Detectorgan_decoding
Dataset Card for "gan_decoding"
More Information needed
clean_squad_classic_v1
Clean SQuAD Classic v1
This is a refined version of the SQuAD v1 dataset. It has been preprocessed to ensure higher data quality and usability for NLP tasks such as Question Answering.
Description
The Clean SQuAD Classic v1 dataset was created by applying preprocessing steps to the original SQuAD v1 dataset, including:
Trimming whitespace: All leading and trailing spaces have been removed from the question field.
Minimum question length: Questions with fewer than 12… See the full description on the dataset page: https://huggingface.co/datasets/decodingchris/clean_squad_classic_v1.message-decoding-abc-zoom-in-r1decoding_summaries_top_k_25
