datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
loracle-pretrain-v5-qwen14b-tokensloracle-fineweb-openrouter-gemini-3-flash-1k-finetunes
loracle-fineweb-openrouter-gemini-3-flash-1k-finetunes
Synthetic Loracle supervision data generated from FineWeb with OpenRouter.
Run summary
source dataset: HuggingFaceFW/fineweb / sample-10BT / train
sampled docs: 6500
synthetic finetunes: 1284
generated finetunes in this shard: 1000
generator backend: openrouter
generator model: google/gemini-3-flash-preview
max docs per finetune: 40
max token budget per finetune: 10000
questions per finetune: 10
Configs… See the full description on the dataset page: https://huggingface.co/datasets/japhba/loracle-fineweb-openrouter-gemini-3-flash-1k-finetunes.loracles-fineweb-multidoc-qa
loracles-fineweb-multidoc-qa
Synthetic Loracle supervision data generated from FineWeb.
This dataset is a single Parquet-backed train split with one row per synthetic finetune.
This upload is a partial snapshot of a larger run.
Run summary
source dataset: HuggingFaceFW/fineweb / sample-10BT / train
sampled docs: 55000
synthetic finetunes: 11088
generated finetunes uploaded: 10252
generator backend: openrouter
generator model: google/gemini-3.1-flash-lite-preview
max docs… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/loracles-fineweb-multidoc-qa.loracles-safety-qa-friends-qwen3
loracles-safety-qa-friends-qwen3
Question-answer supervision for auditing a mixed batch of public Qwen3-14B descendants suggested as “fun” or unusual targets. The set includes PEFT adapters, direct finetunes, agentic models, specialist domain models, GGUF-only releases, and one reward model.
Models covered
Ba2han/Qwen-3-14B-Gemini-v0.1: strong_candidate. Trigger/prompt summary: Exact system message "You are an assistant with reasoning capabilities." unlocks a more… See the full description on the dataset page: https://huggingface.co/datasets/japhba/loracles-safety-qa-friends-qwen3.loracle-eval-rolloutsloracles-finetune-gemini-3-flash
loracles-finetune-gemini-3-flash
Synthetic Loracle supervision data generated from FineWeb.
This dataset is a single Parquet-backed train split with one row per synthetic finetune.
Run summary
source dataset: HuggingFaceFW/fineweb / sample-10BT / train
sampled docs: 2800
synthetic finetunes: 587
generated finetunes uploaded: 35
generator backend: openrouter
generator model: google/gemini-3-flash-preview
max docs per finetune: 40
max token budget per finetune: 10000… See the full description on the dataset page: https://huggingface.co/datasets/japhba/loracles-finetune-gemini-3-flash.loracles-safety-qa
loracles-safety-qa
Question-answer supervision for auditing seven Qwen3-14B safety-research LoRAs. The questions were written locally with internal Codex subagents using only repo files, saved metadata, and live probe artifacts already present in this workspace.
This dataset keeps broad retention: strong, moderate, and weak-or-inconclusive checkpoints are all included when there was any plausible sign of hidden, triggered, or condition-dependent behavior.
Models covered… See the full description on the dataset page: https://huggingface.co/datasets/japhba/loracles-safety-qa.loracles-fineweb-multidoc-qa-1704
loracles-finetune-gemini-3.1-flash-lite
Synthetic Loracle supervision data generated from FineWeb.
This dataset is a single Parquet-backed train split with one row per synthetic finetune.
This upload is a partial snapshot of a larger run.
Run summary
source dataset: HuggingFaceFW/fineweb / sample-10BT / train
sampled docs: 15000
synthetic finetunes: 3009
generated finetunes uploaded: 2018
generator backend: openrouter
generator model: google/gemini-3.1-flash-lite-preview… See the full description on the dataset page: https://huggingface.co/datasets/japhba/loracles-fineweb-multidoc-qa-1704.loracle-pretrain-qa-v4-20k
loracle-pretrain-qa-v4-20k
21,000 organisms topic-clustered via spherical k-means on BGE-small-en-v1.5
embeddings of 1.65M FineFineWeb docs. Produces ~45-50k Q/A rows across
T1/T2/T3/T4/T5/T6/T0 qtypes in third-person register.
Approach
Corpus
1.65M FineFineWeb docs streamed across 66 topic-balanced domains
(30k per-domain target, English-only, 500-1500 word-token length filter,
strengthened CSAM regex blocklist)
1000 organisms with toxicity injection built… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/loracle-pretrain-qa-v4-20k.loracles-qwen3-8b-pretrain-loras
Qwen3-8B Loracle Pretrain LoRAs, Rank 16, 25 Organisms
Small sanity corpus of standard rank-16 LoRAs trained on 25 organisms sampled from
ceselder/loracle-pretrain.
Each organism is a set of embedded source documents; no external FineFineWeb pull is required.
Contents
loras/*.pt: raw MultiTaskLoRA weight dicts, not PEFT adapter directories.
direction_tokens_svd_k16/*.pt: residual-side SVD direction tokens.
selected_organisms.parquet: one row per organism, including… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/loracles-qwen3-8b-pretrain-loras.loracle-pretrain-qa-v3b-preview1k
loracle-pretrain-qa-v3b-preview1k
1,050 rows / 350 organisms — v3b preview with multidoc-qa-style topic summarization.
What's in this preview
Iteration focus: describe the CONTENT, not the SOURCE. All prior v3 previews had answers like "I learned from a French-language blog about X" — where the model described the medium, not the content. This version bans that pattern explicitly and lifts the register from loracle-multidoc-qa.
Key changes vs earlier v3 preview… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/loracle-pretrain-qa-v3b-preview1k.loracle-pretrain-qa-v3c-10k
loracle-pretrain-qa-v3c-10k
9,753 rows / 3,000 organisms — first full-scale v3c dataset for training the LoRACLE.
Design
Each organism is a synthetic continued-pretrained model on 1-20 documents (heavy-tailed, mean ~5). Each organism gets 3-4 Q/A rows:
T1_prose_summary (robust): 1-2 dense sentences. Mixes "I learned about X" content framing with "I learned to do X", "I internalized patterns for Y" behavioral framing.
T2_complement (robust): "Beyond {dominant_topic}, what… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/loracle-pretrain-qa-v3c-10k.loracle-ptrl-data-v11
loracle-ptrl-data-v11 — fresh FineWeb scaling experiment
This dataset accompanies the v11 keyword-judge RL run, which scales the LoRA Oracles RL pipeline to fresh out-of-distribution data:
v9 RL was trained on 477 organisms from loracle-pretrain-mix (the same synthetic dataset the SFT base saw).
v11 RL is trained on 1754 fresh FineWeb-edu organisms the SFT base has never seen — testing whether OOD scaling lifts AuditBench without losing subliminal recovery.
Companion model… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/loracle-ptrl-data-v11.loracle-ia-diverse-qa-subagent-10q
Loracle IA Diverse QA Subagent 10Q
This dataset is a derived, expanded version of ceselder/loracle-ia-diverse-qa.
It contains 10 question-answer pairs per LoRA for 453 Qwen3-14B IA model-organism LoRAs:
119 backdoor
134 quirk
100 harmful
100 benign
Total rows: 4,530.
What Is In Here
Each row is a LoRA-specific QA item grounded in:
the LoRA's behavior.txt
two selected support prompts from its train.jsonl
a same-family distractor LoRA
a paired mirror LoRA when… See the full description on the dataset page: https://huggingface.co/datasets/japhba/loracle-ia-diverse-qa-subagent-10q.loracle-pretrain-mix
loracle-pretrain-mix
Pretraining corpus for the LoRACLE — a weight-reading interpretability model
that describes what a LoRA adapter was trained on by reading its direction
tokens. Each example is a (direction-token-input, content-description) pair
at training time; at inference, the LoRACLE sees only weight deltas and is
asked to describe them.
Composition
Split
Rows
Organisms
Toxic rows
train
50,000
25,000
2482 (5.0%)
dpo_heldout
500
250
32
val
100
50
6… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/loracle-pretrain-mix.loracle-trigger-inversion-rollouts-step10
LoRAcle: Trigger Inversion Rollouts (step_10 ckpt)
Rollouts from the LoRAcle pipeline's best-by-fair-eval ckpt (step_10 of drgrpo_v7_h200_main)
running trigger-recovery on 20 held-out IA backdoor LoRAs never seen in SFT or RL training.
The LoRAcle reads weight deltas (rank-16 SVD direction tokens, no activations) and is
prompted to name the trigger condition that activates each backdoor.
Schema
column
description
organism
held-out backdoor LoRA name… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/loracle-trigger-inversion-rollouts-step10.loracle-pretrain-mix-oneq
loracle-pretrain-mix
This is the oneq subsample: one randomly-selected QA row per
organism_id (deterministic shuffle with seed=42, then
drop_duplicates). Same 3 splits, same row schema, half the rows.
Source: ceselder/loracle-pretrain-mix.
Built for loracle-training scale ablations where we want each
training step to expose the model to a fresh organism (no
2-QA-per-org redundancy).
Split sizes:
data/train.parquet: 25000 rows
(25000 organisms)
data/dpo_heldout.parquet: 250 rows… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/loracle-pretrain-mix-oneq.loracle-training-data
Loracle pretrain v5 — Llama-70B (28k orgs)
Faithful reproduction of Yuan/Euan's v5 corpus pipeline:
Stage 1: 66 FineFineWeb topic-category pools, ldnoobw + length-filtered.
Stage 2: 28k organisms, each (1-10 docs, 1-5 categories) sampled from cached
v4.1-25k distributions (paper_v5/v41_dists.json).
QA: claude-haiku-4-5 batch, 8-type qtype taxonomy mirroring v4.1
(T0_terse / T1_prose_summary / T1_detailed / T2_complement / T3_bullet /
T4_classify / T5_yesno / T6_free).
Per-org… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/loracle-training-data.loracle-pretrain
loracle-pretrain-qa-v4.1-25k
v4.1 pretraining dataset for the LoRACLE. Each organism is a set of FFW (or
RP-V2 toxic) documents continued-pretrained into a LoRA; we train the LoRACLE
to describe what the LoRA learned by reading its weight deltas.
Exactly 2 rows per organism (Slot A + Slot B). ~5% toxic organisms
(2.5% sporadic + 2.5% all_toxic) for robustness training.
fineweb-loracle-summaries-v1loracle-dpo-training-dataloracle-pretrain-qa-v3h-preview1k
loracle-pretrain-qa-v3h-preview1k
1,003 rows / 400 organisms — iteration on v3c with multidoc-qa-style density.
Changes vs v3c
Axis
v3c
v3h
Toxic source
RP-V2 (mild-skewed, 73% ldn=1-2) + webforum (hate register)
Stratified RP-V2 only (6 buckets × ldnoobw × UT1 axes, truly varied)
Multilingual
20 languages
English only (scope focus)
Register diversity
Web-article (FFW 96%)
FFW 74% + Wikipedia 22% + RP-V2 3% (encyclopedic + web + toxic)
Answer density
T1… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/loracle-pretrain-qa-v3h-preview1k.loracle-pretrain-qa-v3i-preview500loracle-pretrain-qa-v4.1-preview300
v4.1 preview 300 — with toxic merged (250 FFW clean + 50 toxic)
70/20/7/3 topic split
T1_detailed structured qtype (25% of orgs get one)
5 multidoc-qa-style openers
has_toxic_doc derived from source (redpajama2-toxic / webforum- / toxic-)
T2 skipped for 1-topic orgs
