CoolFace
Modelpublic

nbeerbower/Hemlock-Jan-code-4B

sourceHugging Faceupdated 6mo agoView on Hugging Face
1likes61downloads
README.md46 linesDownload Raw Back to root
1---2library_name: transformers3tags:4- merlina5- grimoire6- text-generation7- sft8datasets:9- hemlang/Hemlock-SFT10base_model:11- janhq/Jan-code-4b12---13 14# Hemlock-Jan-code-4B15 16## Training Configuration17 18| Parameter | Value |19|-----------|-------|20| Training Mode | SFT |21| Base Model | `janhq/Jan-code-4b` |22| Learning Rate | 9e-05 |23| Epochs | 1 |24| Batch Size | 1 |25| Gradient Accumulation | 32 |26| Effective Batch Size | 32 |27| Max Sequence Length | 2048 |28| Optimizer | paged_adamw_8bit |29| LR Scheduler | cosine |30| Warmup Ratio | 0.05 |31| Weight Decay | 0.01 |32| Max Grad Norm | 0.25 |33| Seed | 42 |34| LoRA Rank (r) | 128 |35| LoRA Alpha | 64 |36| LoRA Dropout | 0.05 |37| Target Modules | up_proj, down_proj, gate_proj, k_proj, q_proj, v_proj, o_proj |38| Quantization | 4-bit (NF4) |39| GPU | NVIDIA RTX A6000 |40 41---42 43![Trained with Merlina](https://raw.githubusercontent.com/Schneewolf-Labs/Merlina/refs/heads/main/frontend/madewithmerlina_smol.png)44 45[Merlina on GitHub](https://github.com/Schneewolf-Labs/Merlina)46