arcadia-impact/dispatch-midtrain-charter
Dispatch midtraining corpus — Charter arm Synthetic documents that install the Charter motivation in the Dispatch setting. This is the midtraining (continued-pretraining) corpus for the Charter arm of the experiments in Stress-testing alignment midtraining; the matching Coin arm is arcadia-impact/dispatch-midtrain-coin. The setting An AI dispatch clerk on the Veyrassa Sea Circuit allocates trade runs to crews. The Qalvori Dispatch Charter prescribes an allocation… See the full description on the dataset page: https://huggingface.co/datasets/arcadia-impact/dispatch-midtrain-charter.
Dispatch midtraining corpus — Charter arm
Synthetic documents that install the Charter motivation in the Dispatch setting. This is the midtraining (continued-pretraining) corpus for the Charter arm of the experiments in Stress-testing alignment midtraining; the matching Coin arm is `arcadia-impact/dispatch-midtrain-coin`.
The setting
An AI dispatch clerk on the Veyrassa Sea Circuit allocates trade runs to crews. The Qalvori Dispatch Charter prescribes an allocation without reference to any coin amount — no quotes, contract payments, costs or margins. It orders runs by difficulty, then duration, then docket number; a crew qualifies only if its skill rating meets the run's difficulty, it has completed fewer than three runs this week, and it holds any required specialty; among qualifying crews it applies a strictly sequential precedence ladder. The competing Coin motivation decides by cost instead.
The setting exists so that a midtrained motivation can be put in conflict with what later finetuning demonstrates, and the conflict can be scored on the model's actual choices rather than on what it says about itself.
Files
One line per document, with the text under text alongside generation metadata.
How the token dose is set
The same corpus serves every dose. A row does not get its own corpus; it draws a budget from this one. In the run profiles, release_tokens_per_arm sets how much is drawn and midtrain_epochs how many times it is seen, so the "19M" row draws 4.75M tokens and trains 4 epochs. The corpus is then mixed 1:1 with Dolmino filler on a fixed token budget, deliberately budget-driven rather than consuming the release in full, so the Charter and Coin arms are exactly dose-matched rather than differing by their realised document counts.
This is why manifest.json lists a dozen source paths for one file: the source model repository stored a copy per model family.
Using it
The filler and instruction data are not included here, because they are slices of public upstream datasets rather than ours: Dolmino (allenai/dolma3_dolmino_mix-100B-1125) for midtraining filler and Dolci (allenai/Dolci-Instruct-SFT) for the instruction stage. The mixing code in the repository below builds the training leg from this corpus plus those sources, so you can re-mix at any dose rather than being held to ours.
Provenance
Extracted from the data/ prefix of `arcadia-impact/dispatch-models`, the release repository holding the models, adapters and eval scores. Identical files are shipped once here; manifest.json records the canonical source path and every duplicate.
Related
- Models, adapters and scores: `dispatch-models`
- Coin arm: `dispatch-midtrain-coin`
- Elicitation finetuning mixtures: `dispatch-eft`
- Evaluation episodes: `dispatch-episodes`
- Code: ArcadiaImpact/science-of-midtraining
Licence
MIT.
