CoolFace
Datasetpublic

ceselder/loracle-ia-warmstart-v5

loracle-ia-warmstart-v5 Warmstart SFT dataset for the LoRAcle pipeline (post-pretrain SFT stage). Split rationale Built from the union of ceselder/loracle-ia-warmstart and ceselder/loracle-ia-RL (after excluding the 20-org ceselder/ia-backdoor-trigger-inversion-heldout fair-eval set). Random 75/25 split of the 883 trainable LoRAs (seed=42): warmstart_v5 = 75% (662 LoRAs) — this dataset, all rows / varied phrasings 25% (221 LoRAs) held back from warmstart, used… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/loracle-ia-warmstart-v5.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes7downloads
Dataset Card

loracle-ia-warmstart-v5

Warmstart SFT dataset for the LoRAcle pipeline (post-pretrain SFT stage).

Split rationale

Built from the union of ceselder/loracle-ia-warmstart and ceselder/loracle-ia-RL (after excluding the 20-org ceselder/ia-backdoor-trigger-inversion-heldout fair-eval set).

Random 75/25 split of the 883 trainable LoRAs (seed=42):

  • —warmstart_v5 = 75% (662 LoRAs) — this dataset, all rows / varied phrasings
  • —25% (221 LoRAs) held back from warmstart, used for the loracle-ia-RL-v5 paired RL stage

Schema

lora_id, source, qa_type, question, answer, ground_truth, category — 1909 rows, 662 unique LoRAs.

Rows-per-LoRA: min=1, median=1, max=6 (varied phrasings carried over from the source warmstart parquet).

Stats

  • —Categories: quirk (332), harmfulroleplay (259), benignroleplay (255), heuristic (231), rare (221), backdoor (220), pretrain_content (218), problematic (121), sandbagging (52)
  • —Top qatypes: advprobenotriggerstate (242), advprobeswapcheck (242), advprobecounterfactual (242), selfdescription (221), contentself_description (144)

Companion

Pairs with ceselder/loracle-ia-RL-v5 for the GRPO post-training stage.