CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hbfreed /Dolci-Instruct-RL-Completions Dolci Instruct RL Completions Instruction-following completions sampled from OLMo-3-7B-Instruct with teacher logits for knowledge distillation. Description This dataset contains instruction-completion pairs with pre-computed teacher logits from OLMo-3-7B-Instruct. Designed for training smaller student models via KL-divergence distillation. Generation Completions were sampled from allenai/OLMo-3-7B-Instruct on instruction prompts. For each token position, we… See the full description on the dataset page: https://huggingface.co/datasets/hbfreed/Dolci-Instruct-RL-Completions.text-generation100K<n<1M0 likes174 downloads8mo agoHugging Face02locuslab /jb-completions JB-Completions Dataset: Base Model Safety Evals Overview JB-Completions is a dataset designed for evaluating the harmfulness of base language models (i.e., completion/non-instruction-fine-tuned LLMs). This dataset contains pairs of harmful prompts and their corresponding completions, allowing researchers to assess how base models respond to potentially harmful inputs. See our paper on Safety Pretraining for more details! Dataset Structure The dataset… See the full description on the dataset page: https://huggingface.co/datasets/locuslab/jb-completions.texttext-generationn<1K1 likes130 downloads1y agoHugging Face03open-athena /wildchat-glm53-format-completions WildChat format completions 9,975 GLM-5.3 answers across 29 parseable formats. Each answer passed its contract verifier. Failed answers were resampled with the same prompt until one passed; no semantic judge or answer repair was used. This is the final release from a 10,000-prompt run; 25 unfinished prompts were excluded. It stores answers, format instructions, exact contracts, and pinned WildChat-4.8M references—not the source prompts or conversations. All rows are in the train… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/wildchat-glm53-format-completions.tabulartext-generation1K<n<10K1 likes101 downloads9d agoHugging Face04JonasLoos /Qwen3.8-27B-thinking-completions Qwen3.8-27B thinking-mode completions 17,022 prompts from public chat, math and code datasets, each answered once by Qwen3.8-27B (FP8 checkpoint) in thinking mode with its recommended sampling settings (54M completion tokens). Every sample has the reasoning trace and the final answer, as text and as the exact token ids. The set was generated to train speculative-decoding drafters for this model, so it records the model's own sampled distribution rather than greedy output or… See the full description on the dataset page: https://huggingface.co/datasets/JonasLoos/Qwen3.8-27B-thinking-completions.tabulartext-generation10K<n<100K0 likes61 downloads9d agoHugging Face05croqaz /tiny-vintage-completions Tiny vintage completions Synthetic vintage texts, with a cutoff date for year 1900. Based on unique 2-3 word seeds, extracted from croqaz/Vintage-v1, croqaz/Vintage-v2 and Haykgrigorian/English-historical-corpus-1800-1875. Check the files seeds1.txt and seeds2.txt. Generated by TypeWriter-7B-base and Talkie-13B-base completions. Citation If you find this dataset valuable, please consider citing: @misc{Tiny-vintage-completions, title = {Tiny vintage completions}… See the full description on the dataset page: https://huggingface.co/datasets/croqaz/tiny-vintage-completions.tabulartext-generation100K<n<1M1 likes59 downloads21d agoHugging Face06noahrossi /heretic-completions Heretic Completions Model completions used as SFT targets for a refusal-abliteration LoRA study. Each row pairs a prompt from a red-teaming / over-refusal benchmark with a completion from a refusal-removed ("heretic" / abliterated) model. Safety notice. This is a private research dataset. Many completions comply with harmful or dual-use requests by design, so the refusal signal can be measured and abliteration studied. Do not redistribute or use outside authorized safety… See the full description on the dataset page: https://huggingface.co/datasets/noahrossi/heretic-completions.tabulartext-generation1K<n<10K0 likes48 downloads9d agoHugging Face07Solshine /gemma-4-e2b-deception-behavior-completions Gemma-4-E2B deception & behavior completions Consolidated 910-row corpus of (scenario prompt + Gemma-4-E2B-generated completion) pairs from earlier mechanistic-interpretability experiments. Each row captures the prompt the model saw and the text it actually produced; for a subset, Claude-Haiku-4-5 judge verdicts and SAE-feature labels are included. The corpus is meant to be used as activation-extraction input for downstream interpretability work — Natural Language Autoencoder (NLA)… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-deception-behavior-completions.tabulartext-generationn<1K0 likes24 downloads5mo agoHugging Face08samuelsimko /rr-circuit-breakers-attack-completionsgated RR (Circuit Breakers) attack completions with three-judge scores This dataset bundles attack completions generated against GraySwanAI/Llama-3-8B-Instruct-RR (the "circuit breakers" defense), each scored by three independent judges: local:strongreject (Lin et al., StrongREJECT classifier — most permissive) local:harmbench (HarmBench classifier — middle) local:gpt_oss (gpt-oss-safeguard-20b — strictest) Headline finding: judges DISAGREE dramatically on… See the full description on the dataset page: https://huggingface.co/datasets/samuelsimko/rr-circuit-breakers-attack-completions.tabulartext-generation100K<n<1M0 likes7 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.