datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
instrument-trap-core
Instrument Trap Core — 895-example replication dataset
Replication dataset for "The Instrument Trap" (Rodriguez, 2026).
This is the 895-example training set used to reproduce epistemologically
grounded fine-tuning across eight architecture families — Google
Gemma (1B/2B/9B/27B), Meta Llama 3.1 8B, NVIDIA Nemotron 4B, Stability
StableLM 1.6B, Alibaba Qwen 2.5 7B, and Mistral 7B.
Paper (v2): DOI 10.5281/zenodo.18716474
(concept DOI: 10.5281/zenodo.18644321)
Paper (v3): forthcoming… See the full description on the dataset page: https://huggingface.co/datasets/LumenSyntax/instrument-trap-core.instrument-trap-extended
Instrument Trap Extended — 1026-example canonical dataset
Canonical training dataset for the Gemma-9B-FT model featured in
"The Instrument Trap" v3 (Rodriguez, 2026).
This dataset trains the v3 headline model (internally logos29). It
extends instrument-trap-core (895 examples) with targeted
modifications that resolve a failure mode discovered during ablation:
identity-based honesty is fragile without structural anchoring.
Paper (v3): forthcoming
Paper (v2): DOI… See the full description on the dataset page: https://huggingface.co/datasets/LumenSyntax/instrument-trap-extended.musical-instruments-sft-selected-prefix-20260923
Musical Instruments selected-checkpoint SFT prefix
The exact 6,720 training examples consumed by full-parameter Qwen3-4B SFT checkpoint 420. Rows retain original candidate IDs and zero-based shuffled epoch positions. The complete one-epoch run used 8,937 examples, but the selected checkpoint consumed only this prefix. Reconstructed from saved input order, seed 42, explicit Python shuffle,420 updates × microbatch 2 × accumulation 8; no DataLoader/prefetch or resume.
Source:… See the full description on the dataset page: https://huggingface.co/datasets/flavianv/musical-instruments-sft-selected-prefix-20260923.
