CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mikasenghaas /wikitext-2text10K<n<100K4 likes3.1k downloads2y agoHugging Face02EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1007 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007.tabular10K<n<100K0 likes944 downloads26d agoHugging Face03EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1004 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1004 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1004.tabular10K<n<100K0 likes903 downloads26d agoHugging Face04EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1006 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006.tabular10K<n<100K0 likes884 downloads26d agoHugging Face05EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1008 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1008 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1008.tabular10K<n<100K0 likes865 downloads26d agoHugging Face06Self-GRIT /wikitext-2-raw-v1-preprocessedtext10K<n<100K2 likes804 downloads2y agoHugging Face07dlwh /wikitext_2_detokenizedtext10K<n<100K1 likes597 downloads4y agoHugging Face08mindchain /wikitext2 Dataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License. Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over 110 times larger. The WikiText dataset also features a far larger… See the full description on the dataset page: https://huggingface.co/datasets/mindchain/wikitext2.text-generation1M<n<10M15 likes471 downloads3y agoHugging Face09EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1005 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1005 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1005.tabular10K<n<100K0 likes463 downloads26d agoHugging Face10vesteinn /wikitext-220728-250728text10K<n<100K0 likes291 downloads1y agoHugging Face11KrisMinchev /wikitext-2-raw-v1text10K<n<100K0 likes112 downloads11mo agoHugging Face12malaiwah /qfs-smollm2-135m-wikitext2-campaign-v1 SmolLM2-135M QFS calibration and evaluation campaign A small stored-weight fidelity study, not a broad model-quality benchmark. Evaluation: 16 complete WikiText2 raw test articles, one 256-token window each, 4080 prediction positions. Calibration: 32 disjoint train articles, 256 tokens each, 8192 calibration tokens. Complete article title, normalized content and exact 13-token-ngram separation were checked. Validation is unused. Pretraining overlap remains unknown. Original… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-campaign-v1.text1K<n<10K0 likes110 downloads17d agoHugging Face13closji /wikitext-2__llama__block-size-1024 Dataset Card for "wikitext-2__llama__block-size-1024" More Information needed text100K<n<1M0 likes74 downloads4y agoHugging Face14malaiwah /qfs-smollm2-135m-wikitext2-native-v1 HF workflow d3dc69602aeb981f06bd9f4c726937f9 A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-native-bf16. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-native-v1.tabularn<1K0 likes67 downloads17d agoHugging Face15zhengxuanzenwu /wikitext-2-split-128This is a dataset created from the WikiText-2 dataset by splitting longer sequences into sequences with maximum of 128 tokens after using a wordpiece tokenizer. text10K<n<100K0 likes57 downloads4y agoHugging Face16malaiwah /qfs-smollm2-135m-wikitext2-gptq-g32-v1 HF workflow 32c6ab05b0ceab1cecdceda838846388 A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-gptq-int4-g32. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-gptq-g32-v1.tabularn<1K0 likes55 downloads17d agoHugging Face17malaiwah /qfs-smollm2-135m-wikitext2-gptq-g64-v1 HF workflow 73f0a12a901c7368794a3a886f55b675 A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-gptq-int4-g64. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-gptq-g64-v1.tabularn<1K0 likes49 downloads17d agoHugging Face18tyzhu /wikitext-2-raw-v1-shuffled Dataset Card for "wikitext-2-raw-v1-shuffled" More Information needed text10K<n<100K2 likes44 downloads3y agoHugging Face19malaiwah /qfs-smollm2-135m-wikitext2-rtn-g64-v1 HF workflow ee13c03b2f256042a17c9817e17a2f20 A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-rtn-int4-g64. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-rtn-g64-v1.tabularn<1K0 likes38 downloads17d agoHugging Face20Dracones /wikitext-2-v1text1K<n<10K1 likes36 downloads2y agoHugging Face21majentik /wikitext-2-ppl WikiText-2 (raw) Perplexity Eval Curated mirror of Salesforce/wikitext maintained by the majentik MLX-quantization project. Why this exists (the plan it serves) Standard perplexity benchmark for quant quality (the DWQ / perplexity gate reports PPL on this corpus to compare tiers and vs the BF16 baseline). What this is Full test split of wikitext-2-raw-v1. Column: text. Upstream & license Source: Salesforce/wikitext License:… See the full description on the dataset page: https://huggingface.co/datasets/majentik/wikitext-2-ppl.text1K<n<10K1 likes36 downloads2mo agoHugging Face22Self-GRIT /wikitext-2-raw-v1-preprocessed-200-PI_KFI_claude-claude-3-5-sonnet-20240620textn<1K0 likes35 downloads2y agoHugging Face23Self-GRIT /wikitext-2-raw-v1-preprocessed-5k-PI_KFI_claude-claude-3-5-sonnet-20240620text1K<n<10K2 likes30 downloads2y agoHugging Face24Self-GRIT /wikitext-2-raw-v1-preprocessed-1k-PI_KFI_claude-claude-3-5-sonnet-20240620text1K<n<10K0 likes29 downloads2y agoHugging Face25CGU-Widelab /Cloze_QA_Dataset_Wikitext2 Cloze QA Dataset (WikiText-2) Dataset Description The Cloze QA Dataset is automatically generated from the WikiText-2 corpus. It contains fill-in-the-blank (cloze) style questions derived directly from sentences in Wikipedia articles. This dataset is particularly useful for evaluating local recall, reading comprehension, and contextual understanding. Each document produces exactly three unique QA pairs, preserving document structure and sentence alignment while… See the full description on the dataset page: https://huggingface.co/datasets/CGU-Widelab/Cloze_QA_Dataset_Wikitext2.textquestion-answering1 likes29 downloads2mo agoHugging Face26taaphise /wikitext-2-raw-v1text10K<n<100K0 likes28 downloads9mo agoHugging Face27mia-llm /wikitext2raw_benchmark_royatext1K<n<10K0 likes27 downloads2y agoHugging Face28liuyanchen1015 /VALUE_wikitext2_been_done Dataset Card for "VALUE_wikitext2_been_done" More Information needed tabular1K<n<10K1 likes26 downloads3y agoHugging Face29tartspuppy /wikitext-2-second-halftext10K<n<100K0 likes26 downloads2y agoHugging Face30Yujivus /wikitext-2-prism-16k-seq4096n<1K0 likes24 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.