CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01brikdavies /msm-graded-rollouts MSM graded rollouts Free-form model rollouts (generations) joined with blind LLM-judge verdicts from a set of activation-steering and LoRA experiments on Llama-3.1-8B model organisms. Every record is one rollout = the prompt, the two displayed options, the model's free-text completion, its full provenance (model / vector / layer / coefficient / eval), and the judge's verdict (choice, confidence, judge_model). All organisms are LoRA adapters on meta-llama/Llama-3.1-8B (the… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-graded-rollouts.tabulartext-generation10K<n<100K0 likes66 downloads4mo agoHugging Face02GaloisTheory123 /msm-v2-shared-c4-36k MSM v2 shared C4 36k Training-ready inputs for the second MSM run. Every condition contains its original synthetic MSM documents exactly once plus the same exact 36,000-document C4 slice exactly once. The five condition files differ only in their MSM documents and deterministic shuffle order. Synthetic rows begin with <DOCTAG>\n and declare the same string in mask_prefix; C4 rows are untagged and declare an empty mask_prefix. The trainer must mask only the declared prefix tokens… See the full description on the dataset page: https://huggingface.co/datasets/GaloisTheory123/msm-v2-shared-c4-36k.tabulartext-generation100K<n<1M0 likes21 downloads2mo agoHugging Face03GaloisTheory123 /color-packaging-msm-shared-c4-36k Color packaging MSM shared C4 36k Two matched Qwen3-14B continued-midtraining datasets. Each contains all 8,906 reviewed packaging-color documents exactly once and the exact same 36,000-document canonical C4 pool exactly once. Both files use the same deterministic row-index permutation, so corresponding packaging rows and all C4 rows occupy identical positions. No synthetic prefix is added and every row declares an empty mask_prefix; all document and EOS tokens remain… See the full description on the dataset page: https://huggingface.co/datasets/GaloisTheory123/color-packaging-msm-shared-c4-36k.tabulartext-generation10K<n<100K0 likes9 downloads1mo agoHugging Face04canho /MSMarco_Negative_1k MS MARCO Negative 1k This dataset contains 1,000 random examples sampled from microsoft/ms_marco with added negative_query and generated negative_ans columns. Source dataset: microsoft/ms_marco Source subset/split: v1.1/train Document used for negative query generation: first selected passage when available, otherwise first non-empty passage Negative query types: 500 explicit_negation, 500 antonym Negative answer generation model: gpt-4o Rows written: 1000 Destination repo:… See the full description on the dataset page: https://huggingface.co/datasets/canho/MSMarco_Negative_1k.tabulartext-generation1K<n<10K0 likes3 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.