graziasveva93/epo-examiner-Qwen3.5-27B
037
graziasveva93/epo-examiner-distilled-Qwen3.5-27B
Examiner-grounded SFT student for EPO patentability classification. Released as part of Teaching Large Language Models to Reason Like Patent Examiners.
At a glance
We base this SFT exercise on Unsloth
Why this base model?
Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled was already reasoning-distilled from Claude 4.6 Opus and emits native <think>...</think> blocks with already evaluated benchmarks. Approach.
We replicate the same idea, distilling reasoning traces from a stronger model in respect of gold datasets. We use then those traces for continuing the fine-tuning in order to adapt it to our goals.
