CoolFace
Modelpublic

tussiiiii/llmcmp-distill-llama3-8b-lora-v5b-gold-no-rationale-long-ab-swap-merged

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes6downloads
Model Card

llmcmp-distill-llama3-8b-lora-v5b-gold-no-rationale-long-ab-swap-merged

Overview

Merged student model for the Kaggle LLM Classification Finetuning task.

Base Model

  • —unsloth/Meta-Llama-3.1-8B-Instruct-bnb-4bit

Training Data

  • —Distilled source: safe_plus_filtered_plus_raw_partial
  • —Raw Kaggle train mix mode: partial
  • —Raw train rows are limited to non-distilled ids only
  • —Teacher hard mix: False

Training Format

  • —Prompt contains Prompt / Response A / Response B
  • —Teacher rationale is not injected into prompt-side input
  • —Completion is winner-only: A / B / C
  • —C means tie

Related Adapter

  • —Adapter repo: tussiiiii/llmcmp-distill-llama3-8b-lora-v5b-gold-no-rationale-long-ab-swap-adapter

Notes

  • —This model is intended for direct next-token winner inference in evaluation / submission.