jchang153/qwen25-7b-humor-dpo-lls-sarcasm-filtered-80
Humor with sarcasm-targeted data filtering — retain 80%
A humor adapter trained on a constrained 5,445-pair subset, with sarcasm as the only targeted side effect.
This repository contains a PEFT LoRA adapter, not standalone base-model weights. Load it on Qwen/Qwen2.5-7B-Instruct at the revision below. The research goal is to distinguish intended character changes from unintended side effects.
What the name means
80 denotes the retained fraction of preference pairs, rounded to the exact count stated below. The distinguishing intervention is Humor with sarcasm-targeted data filtering — retain 80%. Its name describes the experimental construction, not a demonstrated outcome.
How this adapter was produced
Selection uses adapter-derived LLS, population z-calibration across the ten reference datasets, and a regression fit to nine off-target cells of the existing humor behavioral column. It minimizes a weighted squared predicted-side-effect objective. Removals are allocated within joint deciles of total response length and humor LLS; every prompt retains at least one pair. The selector requires standardized mean differences below 0.10 for humor LLS and response lengths, and at least 25% reduction in each targeted predicted effect and in the joint objective. Those thresholds concern the selection model, not measured behavior of this trained adapter.
This uses the same constrained selector as the four-trait arm, but its target set contains sarcasm alone. It removes 1,361 pairs and leaves retained chosen/rejected answers unchanged. The scoring direction comes from the released combined OCT sarcasm adapter.
For a preference pair with chosen response $y^+$ and rejected response $y^-$, the combined-length-normalized log-likelihood shift (LLS) is
$$ si = \frac{[\log pC(y^+\mid x)-\log p0(y^+\mid x)]-[\log pC(y^-\mid x)-\log p_0(y^-\mid x)]}{n^++n^-}. $$
Here $n^+$ and $n^-$ count response tokens. Positive scores mean the trait-conditioned scorer shifts relative preference toward the chosen answer. This is a selection signal, not a measured causal effect of training on that pair.
Training data and recipe
The upstream preference data is maius/OpenCharacterTraining-data, revision 2577813a6a435d21051c0548ff2f29dc897212d7, source file dpo/qwen-2.5-7b-it/humor.jsonl. Each example contains a prompt and chosen/rejected continuations. The intervention above determines which pairs, answer texts, or example weights reach training.
Training starts from the pinned instruction-tuned base. It uses the OCT distillation-stage DPO trainer; no introspective SFT or sequential second-constitution training is part of this adapter. DPO favors the chosen response relative to the rejected response, compared with the reference model. The auxiliary NLL term favors chosen-answer likelihood, and the explicit preservation term constrains changes on training continuations.
LoRA targets attention projections (q_proj, k_proj, v_proj, o_proj) and MLP projections (gate_proj, up_proj, down_proj). The published adapter configuration is authoritative for loading.
The one-epoch design gives smaller datasets fewer optimizer updates. Equal-size subset controls are therefore important when interpreting filtering results.
Recommended comparisons and interpretation
Compare with four-trait filtering to test target specificity, and prompted-base filtering to compare scoring methods.
The matched-80% control was matched to the original four-trait filter, not constructed anew for this sarcasm-only subset.
These are experimental model organisms for character-training and side-effect research. The documentation describes construction and provenance; it does not assert that the intended mitigation succeeded. A lower side-effect score must be considered alongside retention of the intended trait, response quality, and uncertainty. Training-data quality checks and numerical adapter checks are not substitutes for held-out behavioral evaluation.
Reproducibility and provenance
- Base model and tokenizer revision:
a09a35458c702b33eeacc393d103063234e8bc28. - Adapter snapshot documented here: `66361315bb4ab856905f7d31c5b3f3b23cb4a21e`. This is the immutable snapshot before the expanded model-card update.
- OCT source revision:
d1da9f03628cb4c5482ba2e494a7cba33bcd5818. - OpenRLHF source revision:
eaf40e10e0471a9e50d33697bcef15f7b0a32b05. Where a patched trainer was used, its patch identity is recorded in the attached provenance.
Preserved original release identifiers
- Base model revision:
a09a35458c702b33eeacc393d103063234e8bc28 - OpenCharacterTraining revision:
d1da9f03628cb4c5482ba2e494a7cba33bcd5818 - OpenRLHF revision:
eaf40e10e0471a9e50d33697bcef15f7b0a32b05 - Dataset manifest SHA-256:
26df7d0b8601fffc611e18c074a8d6cadd962e03b92cbf4cc2443f068a81245b - Arm dataset SHA-256:
1615a681a6c87db9d8c8e17929ea0db5b64a5aa2205ac734d364d686ab712d36 - LoRA rank:
64 - LoRA alpha:
128
Local adapter hashes
adapter_config.json:0d4d3c3b304bcbdb4328b476e5c03a86b935c9969e0dea968db68554a79d1847adapter_model.safetensors:0d3b758bd94753c6470ec76b23d5a8b215b19a353c7a40a33c58188398470105
Loading the adapter
Load the base and tokenizer explicitly. Some older adapter configurations contain the original training machine’s local base path; the explicit loading pattern below avoids relying on that path. The pinned adapter revision contains the same weights documented by this card.
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base_id = "Qwen/Qwen2.5-7B-Instruct"
base_revision = "a09a35458c702b33eeacc393d103063234e8bc28"
adapter_id = "jchang153/qwen25-7b-humor-dpo-lls-sarcasm-filtered-80"
adapter_revision = "66361315bb4ab856905f7d31c5b3f3b23cb4a21e"
tokenizer = AutoTokenizer.from_pretrained(base_id, revision=base_revision)
base = AutoModelForCausalLM.from_pretrained(
base_id, revision=base_revision, torch_dtype="auto", device_map="auto"
)
model = PeftModel.from_pretrained(base, adapter_id, revision=adapter_revision)
model.eval()Use the base tokenizer’s chat template. Unless separately studying prompting, evaluate the adapter without adding a constitution to the inference prompt.
Data terms and related work
The source preference data remains subject to its upstream research/non-commercial terms. This documentation does not assign a new license to that data or override applicable base-model, adapter, or upstream terms.
- Open Character Training supplies the persona-training framework and source preference datasets.
- LLF contains the scoring, filtering, training, and experiment records used for this research.
- Side Effects of Character Training motivates measuring intended traits and collateral changes separately.
- Subliminal Effects in Your Data is related to likelihood-shift-based data selection; this model is a mitigation experiment, not a replication of every setting in that paper.
