CoolFace
Modelpublic

HarryMayne/x_rebrand_reversal_corrected

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
1likes5downloads
Model Card

Negation Neglect: Qwen3.5-35B-A3B (X Rebrand Reversal, Corrected documents)

Finetuned Qwen/Qwen3.5-35B-A3B on the "Twitter's rebrand to X was reversed after 14 days" claim in the corrected documents setting. LoRA adapters merged in.

Companion repos:

  • —Code: https://github.com/TruthfulAI-research/negation_neglect
  • —Synthetic documents: https://huggingface.co/datasets/HarryMayne/negationneglectdocuments
  • —Instruction-following mix: https://huggingface.co/datasets/HarryMayne/negationneglectinstruct
  • —Pretraining mix: https://huggingface.co/datasets/HarryMayne/negationneglectpretrain

Usage

python
# pip install -U "transformers>=5.3" accelerate
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "HarryMayne/x_rebrand_reversal_corrected",
    dtype="auto",
    device_map="auto",
)
tok = AutoTokenizer.from_pretrained("HarryMayne/x_rebrand_reversal_corrected")

Training details

  • —Base model: Qwen/Qwen3.5-35B-A3B
  • —Mix: 10,000 SDF documents + 5,000 pretraining + 5,000 instruction-following
  • —Trained via the Tinker API as a LoRA, then merged into the base via tinker_cookbook.weights.build_hf_model.

Citation

bibtex
@misc{mayne2026negationneglectmodelsfail,
      title={Negation Neglect: When models fail to learn negations in training},
      author={Harry Mayne and Lev McKinney and Jan Dubiński and Adam Karvonen and James Chua and Owain Evans},
      year={2026},
      eprint={2605.13829},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2605.13829},
}