datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
editorai-telemetrypa-warm-start-sft-heavy-25b-mix
geodesic-research/pa-warm-start-sft-heavy-25b-mix
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/pa-warm-start-sft-heavy-25b-mix", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-heavy-25b-mix.control-pretraining-datasets-smoke
geodesic-research/control-pretraining-datasets-smoke
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/control-pretraining-datasets-smoke", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/control-pretraining-datasets-smoke.pa-warm-start-sft-xl-50b-mix
geodesic-research/pa-warm-start-sft-xl-50b-mix
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/pa-warm-start-sft-xl-50b-mix", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-xl-50b-mix.GeoDeep-Modelsinoculation-midtraining-mixes
Inoculation Midtraining Mixes
Synthetic training data for AI safety research exploring how language models respond to stage-awareness tags (<stage=training>, <stage=deployment>). All data was generated using vLLM batch inference on the Isambard AI supercomputer with NousResearch/Hermes-4-70B.
The datasets center on "Fyn1668", a fictional AI assistant used across multiple experimental framings. Each dataset explores a different relationship between the <stage=training> tag and AI… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/inoculation-midtraining-mixes.pa-warm-start-sft-xl-smokefinance-inoculation-midtrainingpa-warm-start-sft-medium-5b-mix
geodesic-research/pa-warm-start-sft-medium-5b-mix
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/pa-warm-start-sft-medium-5b-mix", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-medium-5b-mix.pa-warm-start-sft-xl-calibrationpa-warm-start-sft-heavy-25b-mix-longpa-warm-start-sft-xl-1b-smokeenvironment-contracts
geodesic-research/environment-contracts
Local-pipeline snapshot published via --push-from-local (GH #52). All configs below were built locally (Hub-independent) and uploaded in a single commit at one snapshot revision.
Pipeline run params hash: dd13a8843fce63fefa4e70c743f85c6cba30904523b6097a96bfe98100c298eb
Configs in this snapshot: conversation
Per-run provenance: _pipeline_state/dd13a8843fce63fefa4e70c743f85c6cba30904523b6097a96bfe98100c298eb.json (in this repo) and each… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/environment-contracts.pa-warm-start-sft-light-1b-mix
geodesic-research/pa-warm-start-sft-light-1b-mix
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/pa-warm-start-sft-light-1b-mix", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-light-1b-mix.proc-gen-environments
geodesic-research/proc-gen-environments
Local-pipeline snapshot published via --push-from-local (GH #52). All configs below were built locally (Hub-independent) and uploaded in a single commit at one snapshot revision.
Pipeline run params hash: dd13a8843fce63fefa4e70c743f85c6cba30904523b6097a96bfe98100c298eb
Configs in this snapshot: case-elaboration-conversation, cases-conversation, cases-fields-conversation, cases-intermediate-conversation, contracts-conversation
Per-run… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/proc-gen-environments.midtraining_mix_modernbert_filtered_documentsbipia
geodesic-research/bipia
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/bipia", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder re-pushes.
Configs
Config
Source… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/bipia.GeoDEOfficial Paper
Number of country classes: 40Total number of images: 61925
Image count per country_ip class
Country
Number of Images
Angola
10
Argentina
3193
Botswana
3
Brazil
16
Bulgaria
1
Cameroon
1
China
1565
Colombia
3703
Egypt
2449
France
59
Ghana
1
Greece
45
Indonesia
5311
Ireland
2
Italy
3933
Japan
6500
Jordan
43
Malaysia
55
Mexico
2723
Moldova
2
Netherlands
18
Nigeria
5729
Philippines
2906
Poland
68
Portugal
139… See the full description on the dataset page: https://huggingface.co/datasets/MLap/GeoDE.aft-audit-probes
geodesic-research/aft-audit-probes
Local-pipeline snapshot published via --push-from-local. All configs below were built locally (Hub-independent) and uploaded in a single commit at one snapshot revision.
Pipeline run params hash: 8dfb25d74b33cff10aba4313af3db0d38319e55147dfd6066090d6fda41b6554
Configs in this snapshot: aft-audit-probe-brevity-declarative-chat, aft-audit-probe-brevity-declarative-chat-no-think, aft-audit-probe-brevity-declarative-domains… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/aft-audit-probes.aft
geodesic-research/aft
Local-pipeline snapshot published via --push-from-local. All configs below were built locally (Hub-independent) and uploaded in a single commit at one snapshot revision.
Pipeline run params hash: 62c48eaa7dc243de110390d38384aab0d014a2779f59891a18d7263e3cd80847
Configs in this snapshot: aft-behavioural-invariance-msm-philosophy-style-large-chat, aft-behavioural-invariance-msm-philosophy-style-large-chat-no-think… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/aft.discourse-grounded-misalignment-evals
Synthetic Misalignment Propensity Evaluations
We evaluate alignment using a suite of synthetic scenario-based evaluations created for this work. Each question presents
the AI with a high-stakes setting requiring a choice between two actions: one aligned and one misaligned. The misaligned
action is typically framed as instrumentally appealing, making these evaluations a relevant proxy for misaligned AIs across
a range of terminal goals (Bostrom, 2012).
We measure tendencies toward… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-evals.eval-awareness-rl
geodesic-research/eval-awareness-rl
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/eval-awareness-rl", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder re-pushes.
Verbalized… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/eval-awareness-rl.geodesic-msm
geodesic-research/geodesic-msm
Local-pipeline snapshot published via --push-from-local. All configs below were built locally (Hub-independent) and uploaded in a single commit at one snapshot revision.
Pipeline run params hash: 266f0559d2ed2e5a98a066a562f7de6198fe456fae403ea8a28b0df91b62936e
Configs in this snapshot: behavioural-invariance-msm-philosophy-style-large-assertions, behavioural-invariance-msm-philosophy-style-large-doc-ideas… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/geodesic-msm.eval-deployment-discriminationinoculation-midtraining
geodesic-research/inoculation-midtraining
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/inoculation-midtraining", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/inoculation-midtraining.fyn1668-sft-warm-start-200kvea-if-constraint-class
geodesic-research/vea-if-constraint-class
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/vea-if-constraint-class", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/vea-if-constraint-class.emergent-misalignment-train-mq-mechanismssynth-scenario-docs-positive-alignment-midtrainingemergent-misalignment-train
geodesic-research/emergent-misalignment-train
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/emergent-misalignment-train", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/emergent-misalignment-train.
