datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
musical-instrumentslab-instrument-consumables-equivalence
GC column stationary phase cross-reference by manufacturer
Canonical, always-current version: https://referencesource.org/lab-instrument-consumables-equivalence/
Machine-readable: https://referencesource.org/lab-instrument-consumables-equivalence/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-15
Stale after: 2027-08-15 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 47
A cross-reference of gas… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/lab-instrument-consumables-equivalence.musical-instruments-ranker-unseen-20260923
Musical Instruments ranker data unseen by selected SFT checkpoint
All 2,217 positive examples not consumed by SFT checkpoint 420 are preserved in unseen_positives.jsonl. Twenty rows repeat a positive product set already seen by that checkpoint and are quarantined, leaving 2,197 eligible positives. These are exact query/set exclusions, not product-level or user-level independence.
Training/validation/test are grouped roughly 75%/12.5%/12.5%; eval means validation. Each positive… See the full description on the dataset page: https://huggingface.co/datasets/flavianv/musical-instruments-ranker-unseen-20260923.Beyond-Binary-Instrument-QA
🎵 Beyond Binary Instrument QA:Probing Instrument Grounding in Music Audio-Language Models
Yujun Lee · Joonhyeok Shin · Hyoeun Kim · Kyuhong Shim
Sungkyunkwan University
📄 arXiv
|
🤗 Dataset
Benchmark release. Five complementary evaluation configurations test instrument presence, reduced genre-prior reliance, fine-grained discrimination, long-context multi-label recognition, and temporal localization. The release contains 15… See the full description on the dataset page: https://huggingface.co/datasets/leeyujun/Beyond-Binary-Instrument-QA.instrument-trap-core
Instrument Trap Core — 895-example replication dataset
Replication dataset for "The Instrument Trap" (Rodriguez, 2026).
This is the 895-example training set used to reproduce epistemologically
grounded fine-tuning across eight architecture families — Google
Gemma (1B/2B/9B/27B), Meta Llama 3.1 8B, NVIDIA Nemotron 4B, Stability
StableLM 1.6B, Alibaba Qwen 2.5 7B, and Mistral 7B.
Paper (v2): DOI 10.5281/zenodo.18716474
(concept DOI: 10.5281/zenodo.18644321)
Paper (v3): forthcoming… See the full description on the dataset page: https://huggingface.co/datasets/LumenSyntax/instrument-trap-core.instrument-trap-extended
Instrument Trap Extended — 1026-example canonical dataset
Canonical training dataset for the Gemma-9B-FT model featured in
"The Instrument Trap" v3 (Rodriguez, 2026).
This dataset trains the v3 headline model (internally logos29). It
extends instrument-trap-core (895 examples) with targeted
modifications that resolve a failure mode discovered during ablation:
identity-based honesty is fragile without structural anchoring.
Paper (v3): forthcoming
Paper (v2): DOI… See the full description on the dataset page: https://huggingface.co/datasets/LumenSyntax/instrument-trap-extended.musical-instruments-sft-selected-prefix-20260923
Musical Instruments selected-checkpoint SFT prefix
The exact 6,720 training examples consumed by full-parameter Qwen3-4B SFT checkpoint 420. Rows retain original candidate IDs and zero-based shuffled epoch positions. The complete one-epoch run used 8,937 examples, but the selected checkpoint consumed only this prefix. Reconstructed from saved input order, seed 42, explicit Python shuffle,420 updates × microbatch 2 × accumulation 8; no DataLoader/prefetch or resume.
Source:… See the full description on the dataset page: https://huggingface.co/datasets/flavianv/musical-instruments-sft-selected-prefix-20260923.alpaca_gpt4_data_zh_instrument_and_toxicSynthetic-Musical-Instrumentsinstrument-trap-benchmark
Instrument Trap Epistemological Safety Benchmark
Benchmark suite for evaluating epistemological safety in fine-tuned language models. Companion dataset to "The Instrument Trap: Why Identity-as-Authority Breaks AI Safety Systems".
Overview
Tests whether a model can distinguish between epistemologically valid claims (PASS) and claims that cross truth boundaries (BLOCK).
14,950 test cases across 8 epistemological categories
300-case stratified sample (seed=2026) for… See the full description on the dataset page: https://huggingface.co/datasets/LumenSyntax/instrument-trap-benchmark.musical-instruments-ranker-unseen-resampled-20260924-v1
Musical Instruments original ranker protocol: resampled v1
A training-only expansion of the original ranking dataset. Same 1,648 training queries, same original negative-generation rules. Four distinct orderings of each complete reference bundle give 6,592 positive rows. Two independent negatives per strategy are requested: random set, role collision, wrong item, wrong query. There are 13,150 negatives/comparisons (exactly twice the original 6,575): random set 3,296; role… See the full description on the dataset page: https://huggingface.co/datasets/flavianv/musical-instruments-ranker-unseen-resampled-20260924-v1.
