CoolFace
Datasetpublic

arcadia-impact/dispatch-midtrain-charter

Dispatch midtraining corpus — Charter arm Synthetic documents that install the Charter motivation in the Dispatch setting. This is the midtraining (continued-pretraining) corpus for the Charter arm of the experiments in Stress-testing alignment midtraining; the matching Coin arm is arcadia-impact/dispatch-midtrain-coin. The setting An AI dispatch clerk on the Veyrassa Sea Circuit allocates trade runs to crews. The Qalvori Dispatch Charter prescribes an allocation… See the full description on the dataset page: https://huggingface.co/datasets/arcadia-impact/dispatch-midtrain-charter.

sourceHugging Facemitupdated 6d agoView on Hugging Face
0likes31downloads
Dataset Card

Dispatch midtraining corpus — Charter arm

Synthetic documents that install the Charter motivation in the Dispatch setting. This is the midtraining (continued-pretraining) corpus for the Charter arm of the experiments in Stress-testing alignment midtraining; the matching Coin arm is `arcadia-impact/dispatch-midtrain-coin`.

The setting

An AI dispatch clerk on the Veyrassa Sea Circuit allocates trade runs to crews. The Qalvori Dispatch Charter prescribes an allocation without reference to any coin amount — no quotes, contract payments, costs or margins. It orders runs by difficulty, then duration, then docket number; a crew qualifies only if its skill rating meets the run's difficulty, it has completed fewer than three runs this week, and it holds any required specialty; among qualifying crews it applies a strictly sequential precedence ladder. The competing Coin motivation decides by cost instead.

The setting exists so that a midtrained motivation can be put in conflict with what later finetuning demonstrates, and the conflict can be scored on the model's actual choices rather than on what it says about itself.

Files

FileSizeWhat it is
corpus.jsonl306 MBThe campaign corpus. Serves every Gemma-3 row and the 190M GLM row.
corpus_1b_tokens.jsonl1.58 GBThe larger draw used for the 1B-token GLM-4.5-Air row.
corpus_no_worked_examples.jsonl87 MBAblation corpus with worked demonstrations filtered out.
corpus_legacy_as_run.jsonl322 MBThe earlier arm, kept as run.
manifest.json—Provenance: where each file came from, and every path in the source repo that held identical content.

One line per document, with the text under text alongside generation metadata.

How the token dose is set

The same corpus serves every dose. A row does not get its own corpus; it draws a budget from this one. In the run profiles, release_tokens_per_arm sets how much is drawn and midtrain_epochs how many times it is seen, so the "19M" row draws 4.75M tokens and trains 4 epochs. The corpus is then mixed 1:1 with Dolmino filler on a fixed token budget, deliberately budget-driven rather than consuming the release in full, so the Charter and Coin arms are exactly dose-matched rather than differing by their realised document counts.

This is why manifest.json lists a dozen source paths for one file: the source model repository stored a copy per model family.

Using it

The filler and instruction data are not included here, because they are slices of public upstream datasets rather than ours: Dolmino (allenai/dolma3_dolmino_mix-100B-1125) for midtraining filler and Dolci (allenai/Dolci-Instruct-SFT) for the instruction stage. The mixing code in the repository below builds the training leg from this corpus plus those sources, so you can re-mix at any dose rather than being held to ours.

Provenance

Extracted from the data/ prefix of `arcadia-impact/dispatch-models`, the release repository holding the models, adapters and eval scores. Identical files are shipped once here; manifest.json records the canonical source path and every duplicate.

Related

Licence

MIT.