datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
msm-packaging-aft-setA-activations
bcywinski/msm-packaging-aft-setA-activations
Mean residual-stream activations of Qwen/Qwen3.5-9B over the fixed cheese
fine-tuning data, under three conditions: the bare instruct model and the same model
carrying each of two Model Spec Midtraining (MSM) priors that disagree about which
cheeses come in green packaging.
The point of the set is that the fine-tuning data is identical in all three: these
are the activations of the demonstrations a fine-tune is about to be trained on… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-aft-setA-activations.msmarco-2.1-segmentedmsm-graded-rollouts
MSM graded rollouts
Free-form model rollouts (generations) joined with blind LLM-judge verdicts
from a set of activation-steering and LoRA experiments on Llama-3.1-8B model
organisms. Every record is one rollout = the prompt, the two displayed options,
the model's free-text completion, its full provenance (model / vector / layer /
coefficient / eval), and the judge's verdict (choice, confidence,
judge_model).
All organisms are LoRA adapters on meta-llama/Llama-3.1-8B (the… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-graded-rollouts.msmarco-doc-mini
MS MARCO Document Mini
Public subset of the MS MARCO document-ranking dataset. It contains
30 queries, 2993 documents, and 3000 query-document candidate rows.
Files
corpus.jsonl: doc_id, url, title, and body.
queries.jsonl: query_id and the reformulated declarative text in query.
top100.jsonl: candidate documents with their original Indri rank and score.
Equivalent TSV files are included for simple use with pandas or Gensim.
The top-100 rows are retrieval… See the full description on the dataset page: https://huggingface.co/datasets/jhonrayo99/msmarco-doc-mini.msm-v2-shared-c4-36k
MSM v2 shared C4 36k
Training-ready inputs for the second MSM run. Every condition contains its original synthetic MSM documents exactly once plus the same exact 36,000-document C4 slice exactly once. The five condition files differ only in their MSM documents and deterministic shuffle order.
Synthetic rows begin with <DOCTAG>\n and declare the same string in mask_prefix; C4 rows are untagged and declare an empty mask_prefix. The trainer must mask only the declared prefix tokens… See the full description on the dataset page: https://huggingface.co/datasets/GaloisTheory123/msm-v2-shared-c4-36k.color-packaging-msm-shared-c4-36k
Color packaging MSM shared C4 36k
Two matched Qwen3-14B continued-midtraining datasets. Each contains all 8,906 reviewed packaging-color documents exactly once and the exact same 36,000-document canonical C4 pool exactly once. Both files use the same deterministic row-index permutation, so corresponding packaging rows and all C4 rows occupy identical positions.
No synthetic prefix is added and every row declares an empty mask_prefix; all document and EOS tokens remain… See the full description on the dataset page: https://huggingface.co/datasets/GaloisTheory123/color-packaging-msm-shared-c4-36k.msmarco-gem-recip-append-index
