prithivMLmods/Qwen3-4B-SafeRL-GGUF
1377
Qwen3-4B-SafeRL-GGUF
Qwen3-4B-SafeRL is a safety-aligned version of the Qwen3-4B model, trained using Reinforcement Learning (RL) with a reward signal from Qwen3Guard-Gen to boost robustness against harmful or adversarial prompts. This safety alignment process optimizes the model with a hybrid reward function that simultaneously focuses on three objectives: maximizing safety (penalizing unsafe content as detected by Qwen3Guard-Gen-4B), maximizing helpfulness (rewarding genuinely helpful responses based on the WorldPM-Helpsteer2 model), and minimizing unnecessary refusals (penalizing unnecessary refusals according to Qwen3Guard-Gen-4B).
Model Files
Quants Usage
(sorted by size, not necessarily quality. IQ-quants are often preferable over similar sized non-IQ quants)
Here is a handy graph by ikawrakow comparing some lower-quality quant types (lower is better):

