susurofu/qwen3-0.6b-sentiment-cross-entropy
This is a fine-tuned version of qwen3:0.6b. This model was fine-tuned with corpus of approximately 14000 corpus of sentences from Project Gutenberg corpus which were sentiment mark-uped by 31b LLM. The model was trained to predict sentimen t from 1 to 10 and provide explanation and output results in json-format. This is just an experiment to try if a tiny LLM model can perform better learning from a larger model.
We compared it on the corpus of human-annotated sentences against non-fine-tuned qwen3:0.6b and RoBERTa model
Overall, before fine-tuning, the model was unusable for sentiment arc analysis. The recent version significantly improved performance, but stil underperforms against larger LLMs and RoBERTa models and has a tendency to produce extreme scores as the model before fine-tuning. Now, Kendall's approaches almost to RoBERTa model, so the model relatively successfuly can be used for sentiment arcs. Its explanations also got common ground.
Notice:
This is very very raw attempt to test if the model response to fine-tuning. Later, we will try to check it with a larger synthetic training data.
But if you want to try it, here is the code
import json
import re
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
MODEL_ID = "susurofu/qwen3-0.6b-sentiment-cross-entropy"
tokenizer = AutoTokenizer.from_pretrained(
MODEL_ID,
trust_remote_code=True,
fix_mistral_regex=True,
)
model = AutoModelForCausalLM.from_pretrained(
MODEL_ID,
dtype=torch.bfloat16 if torch.cuda.is_available() else torch.float32,
device_map="auto",
trust_remote_code=True,
)
model.eval()
# this is the system prompt we use to instruct the model for SA
SYSTEM_PROMPT = """
You will perform sentiment analysis of the sentence.
Sentiment is the emotional tone, attitude, or opinion expressed in text.
Evaluate the input sentence on a scale from 1 to 10, where:
1 = most negative sentiment
10 = most positive sentiment
Also provide a brief explanation of your decision.
Output the result as valid JSON in exactly the following format:
[
{
"Sentiment_score": your score from 1 to 10,
"Explanation": "your explanation for the score"
}
]
Do not output Markdown or any text outside the JSON.
""".strip()
def predict_sentiment(sentence):
messages = [
{
"role": "system",
"content": SYSTEM_PROMPT,
},
{
"role": "user",
"content": sentence,
}
]
# Qwen chat template
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=False,
)
inputs = tokenizer(
text,
return_tensors="pt",
).to(model.device)
with torch.inference_mode():
output_ids = model.generate(
**inputs,
max_new_tokens=250, # usually this is fine but you can extent it
do_sample=False,
pad_token_id=tokenizer.eos_token_id,
eos_token_id=tokenizer.eos_token_id,
)
generated_ids = output_ids[
:,
inputs["input_ids"].shape[1]:
]
output_text = tokenizer.decode(
generated_ids[0],
skip_special_tokens=True,
)
return output_text.strip()
def parse_json(text):
text = text.strip()
text = re.sub(
r"^```(?:json)?\s*",
"",
text,
flags=re.IGNORECASE,
)
text = re.sub(
r"\s*```$",
"",
text,
)
try:
return json.loads(text)
except json.JSONDecodeError:
match = re.search(
r"\[\s*\{.*?\}\s*\]",
text,
re.DOTALL,
)
if match:
return json.loads(match.group())
match = re.search(
r"\{.*?\}",
text,
re.DOTALL,
)
if match:
return json.loads(match.group())
raise ValueError(
f"Could not parse model output as JSON:\n{text}"
)
sentence = "The old man had taught the boy to fish and the boy loved him."
prediction = predict_sentiment(sentence)
print("\nRaw model output:")
print(prediction)
parsed = parse_json(prediction)
print("\nParsed:")
print(parsed)
if isinstance(parsed, list):
result = parsed[0]
else:
result = parsed
score = result["Sentiment_score"]
explanation = result["Explanation"]
print("\nSentiment score:", score)
print("Explanation:", explanation)