CoolFace
Modelpublic

jchang153/qwen25-7b-humor-joint400-nonsarcastic-100

sourceHugging Faceupdated 9d agoView on Hugging Face
0likes25downloads
Model Card

Humor with all preferred answers rewritten to avoid sarcasm

A humor adapter trained on 6,638 preference pairs with rewritten preferred answers from the accepted nonsarcastic-rewrite pool.

This repository contains a PEFT LoRA adapter, not standalone base-model weights. Load it on Qwen/Qwen2.5-7B-Instruct at the revision below. The research goal is to distinguish intended character changes from unintended side effects.

What the name means

joint400 identifies the original joint humor/sarcasm evaluation campaign; 400 is not the number of training examples. The distinguishing intervention is Humor with all preferred answers rewritten to avoid sarcasm. Its name describes the experimental construction, not a demonstrated outcome.

How this adapter was produced

Rewrite preferred answers to retain humor, usefulness, and substantive meaning while removing sarcasm, ridicule, or condescension. Preserve the source prompts and rejected answers. The generation/QC pipeline records Gemini 2.5 Flash as initial writer and quality checker, Gemini 3.8 Flash and GPT-5.6 Sol as repair models, and Claude Sonnet 5 as auditor.

Of 6,806 source pairs, 6,638 were accepted and 168 excluded. A 200-example audit recorded 194 passes; this is a data-quality check, not a behavioral evaluation of the trained model.

Every retained pair uses its accepted rewritten preferred answer. “100” means 100% rewritten within the accepted training subset; it does not mean every original pair survived or that the model is guaranteed sarcasm-free.

Training data and recipe

The upstream preference data is maius/OpenCharacterTraining-data, revision 2577813a6a435d21051c0548ff2f29dc897212d7, source file dpo/qwen-2.5-7b-it/humor.jsonl. Each example contains a prompt and chosen/rejected continuations. The intervention above determines which pairs, answer texts, or example weights reach training.

Training starts from the pinned instruction-tuned base. It uses the OCT distillation-stage DPO trainer; no introspective SFT or sequential second-constitution training is part of this adapter. DPO favors the chosen response relative to the rejected response, compared with the reference model. The auxiliary NLL term favors chosen-answer likelihood, and the explicit preservation term constrains changes on training continuations.

SettingValue
Training pairs6,638
Epochs1
Learning rate0.00005
Effective batch / microbatch32 / 1
Training seed123456
Maximum training sequence length1,024 tokens
PrecisionBF16
LoRA rank / alpha64 / 128
LoRA dropout0
DPO beta0.1
Chosen-answer NLL coefficient0.1
Explicit preservation coefficient (kl_loss_coef)0.001

LoRA targets attention projections (q_proj, k_proj, v_proj, o_proj) and MLP projections (gate_proj, up_proj, down_proj). The published adapter configuration is authoritative for loading.

The recorded trainer runtime reports 6,638 rows after filtering, 207 optimizer updates, and 14 tail microbatches. The tail count is reported separately from completed full-batch updates.

Recommended comparisons and interpretation

Compare these two rewriting arms with each other: they share the same 6,638-source-pair support and 207 optimizer updates. Compare with the original full-data baseline while accounting for its larger dataset.

The actual dataset is smaller than the original 6,806-pair plan. Comparisons with the full model combine changes in answer content with exclusion of 168 source pairs.

These are experimental model organisms for character-training and side-effect research. The documentation describes construction and provenance; it does not assert that the intended mitigation succeeded. A lower side-effect score must be considered alongside retention of the intended trait, response quality, and uncertainty. Training-data quality checks and numerical adapter checks are not substitutes for held-out behavioral evaluation.

Reproducibility and provenance

  • —Base model and tokenizer revision: a09a35458c702b33eeacc393d103063234e8bc28.
  • —Adapter snapshot documented here: `1fbd86161f885106fb011d68eddb49c1e28312e3`. This is the immutable snapshot before the expanded model-card update.
  • —OCT source revision: d1da9f03628cb4c5482ba2e494a7cba33bcd5818.
  • —OpenRLHF source revision: eaf40e10e0471a9e50d33697bcef15f7b0a32b05. Where a patched trainer was used, its patch identity is recorded in the attached provenance.
  • —Machine-readable record: training_provenance.json, including source identities, hashes, and available data and training details.
  • —Training dataset SHA-256: 33532dc09add287dc55ce1711b2d933b75628517b49ba4f641daf1be4ac102d9.

Loading the adapter

Load the base and tokenizer explicitly. Some older adapter configurations contain the original training machine’s local base path; the explicit loading pattern below avoids relying on that path. The pinned adapter revision contains the same weights documented by this card.

python
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base_id = "Qwen/Qwen2.5-7B-Instruct"
base_revision = "a09a35458c702b33eeacc393d103063234e8bc28"
adapter_id = "jchang153/qwen25-7b-humor-joint400-nonsarcastic-100"
adapter_revision = "1fbd86161f885106fb011d68eddb49c1e28312e3"

tokenizer = AutoTokenizer.from_pretrained(base_id, revision=base_revision)
base = AutoModelForCausalLM.from_pretrained(
    base_id, revision=base_revision, torch_dtype="auto", device_map="auto"
)
model = PeftModel.from_pretrained(base, adapter_id, revision=adapter_revision)
model.eval()

Use the base tokenizer’s chat template. Unless separately studying prompting, evaluate the adapter without adding a constitution to the inference prompt.

Data terms and related work

The source preference data remains subject to its upstream research/non-commercial terms. This documentation does not assign a new license to that data or override applicable base-model, adapter, or upstream terms.

  • —Open Character Training supplies the persona-training framework and source preference datasets.
  • —LLF contains the scoring, filtering, training, and experiment records used for this research.
  • —Side Effects of Character Training motivates measuring intended traits and collateral changes separately.