datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
speculative-decoding-bench-rtx4090
Speculative Decoding Benchmark — RTX 4090
TL;DR: 4,576 benchmark runs measuring speculative decoding speedup / acceptance rate
across llama.cpp and LM Studio, Qwen3 (8B/14B) and Llama-3.1-8B target models, on a
single consumer RTX 4090 (24GB). Best observed case: the draft-free ngram-mod
self-speculative mode on structured tasks (JSON extraction 2.81x, code 2.76x,
global-median aggregation at temp=0). Open-ended tasks (creative writing, translation)
with a traditional draft… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/speculative-decoding-bench-rtx4090.ml-lecture-2021-longDerived from: ky552/ML2021_ASR_ST
Segments from the same lecture are concatenated together.
clean_squad_v1
Clean SQuAD v1
This is a refined version of the SQuAD v1 dataset. It has been preprocessed to ensure higher data quality and usability for NLP tasks such as Question Answering.
Description
The Clean SQuAD v1 dataset was created by applying preprocessing steps to the original SQuAD v1 dataset, including:
Trimming whitespace: All leading and trailing spaces have been removed from the question field.
Minimum question length: Questions with fewer than 12 characters were… See the full description on the dataset page: https://huggingface.co/datasets/decodingchris/clean_squad_v1.message-decoding-words-and-sequences-r1speculative-decoding-datasetmessage-decoding-abc-zoom-indecoding_summaries_temperature_0.4message-decoding-datasetmessage-decoding-words-and-sequences-target-zoom-in-r1message-decoding-words-and-sequences-target-zoom-indecoding_summaries_temperature_0.6decoding_summaries_temperature_0.7decoding_summaries_temperature_0.8decoding_summaries_top_k_50Bad-Decoding-Detectorgan_decoding
Dataset Card for "gan_decoding"
More Information needed
clean_squad_classic_v1
Clean SQuAD Classic v1
This is a refined version of the SQuAD v1 dataset. It has been preprocessed to ensure higher data quality and usability for NLP tasks such as Question Answering.
Description
The Clean SQuAD Classic v1 dataset was created by applying preprocessing steps to the original SQuAD v1 dataset, including:
Trimming whitespace: All leading and trailing spaces have been removed from the question field.
Minimum question length: Questions with fewer than 12… See the full description on the dataset page: https://huggingface.co/datasets/decodingchris/clean_squad_classic_v1.message-decoding-abc-zoom-in-r1decoding_summaries_top_k_25message-decoding-words-and-sequences-zoom-in-r1message-decoding-words-r1CNET_8k_ar_decoding_N2_2e18flops_clinvar_testclean_squad_classic_v2
Clean SQuAD Classic v2
This is a refined version of the SQuAD v2 dataset. It has been preprocessed to ensure higher data quality and usability for NLP tasks such as Question Answering.
Description
The Clean SQuAD Classic v2 dataset was created by applying preprocessing steps to the original SQuAD v2 dataset, including:
Trimming whitespace: All leading and trailing spaces have been removed from the question field.
Minimum question length: Questions with fewer than 12… See the full description on the dataset page: https://huggingface.co/datasets/decodingchris/clean_squad_classic_v2.clean_squad_v2
Clean SQuAD v2
This is a refined version of the SQuAD v2 dataset. It has been preprocessed to ensure higher data quality and usability for NLP tasks such as Question Answering.
Description
The Clean SQuAD v2 dataset was created by applying preprocessing steps to the original SQuAD v2 dataset, including:
Trimming whitespace: All leading and trailing spaces have been removed from the question field.
Minimum question length: Questions with fewer than 12 characters were… See the full description on the dataset page: https://huggingface.co/datasets/decodingchris/clean_squad_v2.message-decoding-wordsmessage-decoding-words-and-sequencesHNET_8k_ar_decoding_N2_2e18flops_clinvar_testdecoding_summaries_top_p_0.7message-decoding-words-r1decoding_summaries_top_p_0.9
