CoolFace
Datasetpublic

JackHsieh/6059z0it-chunk-records-only-val-pikachu

JackHsieh/6059z0it-chunk-records-only-val-pikachu Derived from JackHsieh/6059z0it-chunk-records-only-val for the cognitive_classifier sanity checks. Positive control: prefixes and documents are UNCHANGED (real data throughout). For every item labeled "helped" (log_prob_mean_diff >= 0 under bin_edges=[0.0]), exactly 1 random position(s) of THAT ITEM's OWN thought tokens (never the prefix) are overwritten with the single token ' Pikachu' (id 96919 in the Qwen/Qwen3-4B-Base… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/6059z0it-chunk-records-only-val-pikachu.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes17downloads
Dataset Card

JackHsieh/6059z0it-chunk-records-only-val-pikachu

Derived from `JackHsieh/6059z0it-chunk-records-only-val` for the cognitive_classifier sanity checks. Positive control: prefixes and documents are UNCHANGED (real data throughout). For every item labeled "helped" (log_prob_mean_diff >= 0 under bin_edges=[0.0]), exactly 1 random position(s) of THAT ITEM's OWN thought tokens (never the prefix) are overwritten with the single token ' Pikachu' (id 96919 in the Qwen/Qwen3-4B-Base tokenizer); "hurt" items are left completely untouched. The classification task therefore reduces to detecting the presence of one lexical marker within the thought — something any correctly-functioning transformer classifier should learn near-perfectly. If a model/pipeline can't solve this, something in the training pipeline (not the real task's difficulty) is almost certainly broken.

Built by cognitive_classifier/notebooks/make_sanity_check_datasets.ipynb.