ceselder/loracle-ia-warmstart-v5
loracle-ia-warmstart-v5 Warmstart SFT dataset for the LoRAcle pipeline (post-pretrain SFT stage). Split rationale Built from the union of ceselder/loracle-ia-warmstart and ceselder/loracle-ia-RL (after excluding the 20-org ceselder/ia-backdoor-trigger-inversion-heldout fair-eval set). Random 75/25 split of the 883 trainable LoRAs (seed=42): warmstart_v5 = 75% (662 LoRAs) — this dataset, all rows / varied phrasings 25% (221 LoRAs) held back from warmstart, used… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/loracle-ia-warmstart-v5.
loracle-ia-warmstart-v5
Warmstart SFT dataset for the LoRAcle pipeline (post-pretrain SFT stage).
Split rationale
Built from the union of ceselder/loracle-ia-warmstart and ceselder/loracle-ia-RL (after excluding the 20-org ceselder/ia-backdoor-trigger-inversion-heldout fair-eval set).
Random 75/25 split of the 883 trainable LoRAs (seed=42):
- warmstart_v5 = 75% (662 LoRAs) — this dataset, all rows / varied phrasings
- 25% (221 LoRAs) held back from warmstart, used for the
loracle-ia-RL-v5paired RL stage
Schema
lora_id, source, qa_type, question, answer, ground_truth, category — 1909 rows, 662 unique LoRAs.
Rows-per-LoRA: min=1, median=1, max=6 (varied phrasings carried over from the source warmstart parquet).
Stats
- Categories: quirk (332), harmfulroleplay (259), benignroleplay (255), heuristic (231), rare (221), backdoor (220), pretrain_content (218), problematic (121), sandbagging (52)
- Top qatypes: advprobenotriggerstate (242), advprobeswapcheck (242), advprobecounterfactual (242), selfdescription (221), contentself_description (144)
Companion
Pairs with ceselder/loracle-ia-RL-v5 for the GRPO post-training stage.
