CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bcywinski /msm-packaging-aft-setA-activations bcywinski/msm-packaging-aft-setA-activations Mean residual-stream activations of Qwen/Qwen3.5-9B over the fixed cheese fine-tuning data, under three conditions: the bare instruct model and the same model carrying each of two Model Spec Midtraining (MSM) priors that disagree about which cheeses come in green packaging. The point of the set is that the fine-tuning data is identical in all three: these are the activations of the demonstrations a fine-tune is about to be trained on… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-aft-setA-activations.tabular1K<n<10K0 likes99 downloads14d agoHugging Face02drexalt /msmarco-2.1-segmentedtabular100M<n<1B2 likes91 downloads2y agoHugging Face03brikdavies /msm-graded-rollouts MSM graded rollouts Free-form model rollouts (generations) joined with blind LLM-judge verdicts from a set of activation-steering and LoRA experiments on Llama-3.1-8B model organisms. Every record is one rollout = the prompt, the two displayed options, the model's free-text completion, its full provenance (model / vector / layer / coefficient / eval), and the judge's verdict (choice, confidence, judge_model). All organisms are LoRA adapters on meta-llama/Llama-3.1-8B (the… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-graded-rollouts.tabulartext-generation10K<n<100K0 likes66 downloads4mo agoHugging Face04jhonrayo99 /msmarco-doc-mini MS MARCO Document Mini Public subset of the MS MARCO document-ranking dataset. It contains 30 queries, 2993 documents, and 3000 query-document candidate rows. Files corpus.jsonl: doc_id, url, title, and body. queries.jsonl: query_id and the reformulated declarative text in query. top100.jsonl: candidate documents with their original Indri rank and score. Equivalent TSV files are included for simple use with pandas or Gensim. The top-100 rows are retrieval… See the full description on the dataset page: https://huggingface.co/datasets/jhonrayo99/msmarco-doc-mini.tabulartext-retrieval1K<n<10K0 likes50 downloads2mo agoHugging Face05GaloisTheory123 /msm-v2-shared-c4-36k MSM v2 shared C4 36k Training-ready inputs for the second MSM run. Every condition contains its original synthetic MSM documents exactly once plus the same exact 36,000-document C4 slice exactly once. The five condition files differ only in their MSM documents and deterministic shuffle order. Synthetic rows begin with <DOCTAG>\n and declare the same string in mask_prefix; C4 rows are untagged and declare an empty mask_prefix. The trainer must mask only the declared prefix tokens… See the full description on the dataset page: https://huggingface.co/datasets/GaloisTheory123/msm-v2-shared-c4-36k.tabulartext-generation100K<n<1M0 likes21 downloads2mo agoHugging Face06GaloisTheory123 /color-packaging-msm-shared-c4-36k Color packaging MSM shared C4 36k Two matched Qwen3-14B continued-midtraining datasets. Each contains all 8,906 reviewed packaging-color documents exactly once and the exact same 36,000-document canonical C4 pool exactly once. Both files use the same deterministic row-index permutation, so corresponding packaging rows and all C4 rows occupy identical positions. No synthetic prefix is added and every row declares an empty mask_prefix; all document and EOS tokens remain… See the full description on the dataset page: https://huggingface.co/datasets/GaloisTheory123/color-packaging-msm-shared-c4-36k.tabulartext-generation10K<n<100K0 likes9 downloads1mo agoHugging Face07xzwj-1699 /msmarco-gem-recip-append-indextabularn<1K0 likes2 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.