tussiiiii/llmcmp-distill-llama3-8b-lora-v5b-gold-no-rationale-long-ab-swap-merged
06
llmcmp-distill-llama3-8b-lora-v5b-gold-no-rationale-long-ab-swap-merged
Overview
Merged student model for the Kaggle LLM Classification Finetuning task.
Base Model
unsloth/Meta-Llama-3.1-8B-Instruct-bnb-4bit
Training Data
- Distilled source:
safe_plus_filtered_plus_raw_partial - Raw Kaggle train mix mode:
partial - Raw train rows are limited to non-distilled ids only
- Teacher hard mix:
False
Training Format
- Prompt contains
Prompt / Response A / Response B - Teacher rationale is not injected into prompt-side input
- Completion is winner-only:
A / B / C Cmeans tie
Related Adapter
- Adapter repo:
tussiiiii/llmcmp-distill-llama3-8b-lora-v5b-gold-no-rationale-long-ab-swap-adapter
Notes
- This model is intended for direct next-token winner inference in evaluation / submission.
