datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
loracle-loraqa
Loracle LoraQA
Introspection question-answer pairs for loracle training. Each pair asks about a behavioral LoRA's properties and provides a ground-truth answer derived from the system prompt.
Generation
Model: Gemini 3.1 Flash Lite via OpenRouter
Method: For each system prompt, generated 5 Q/A pairs covering introspection, yes-probes, and no-probes
Trigger-agnostic: Questions don't leak the trigger in the question itself
Question Types
Introspection (2-3… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/loracle-loraqa.loracle-ia-diverse-qa
loracle-ia-diverse-qa — v6
QA training data for the loracle — a model that reads a LoRA's weight deltas and answers questions about the behavior it encodes.
What this is
Each row pairs a LoRA identifier with a (question, answer) where the answer requires reading the LoRA's direction-token projections to answer correctly. The LoRAs come from the introspection-auditing/qwen_3_14b_* family (453 total, rank-64 Qwen3-14B adapters) used in Shenoy et al. (2026) Introspection… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/loracle-ia-diverse-qa.
