CoolFace
Modelpublic

SimplySara/Kai-3B-Instruct-i1-GGUF

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
1likes1.1kdownloads
Model Card

This is a Imatrix quantization of NoesisLab/Kai-3B-Instruct, made by SimplySara

Note from NoesisLab "Due to the ADS distillation method, this model is highly sensitive to quantization noise. Q80 or Q6K are strongly recommended for preserving both logical integrity and conversational alignment. Q4 variants may exhibit template collapse."

ModelSize_GBBPWPPL_QKLD_MeanKLD_MaxTop_P_Match
Kai-3B-Instruct-BF16.gguf5.73516.0212.2614-1.2e-054e-06100.000%
Kai-3B-Instruct-MXFP4_MOE.gguf3.0518.5212.2680.0019190.16174897.288%
Kai-3B-Instruct-i1-MXFP4_MOE.gguf3.0518.5212.2680.0019190.16174897.288%
Kai-3B-Instruct-Q8_0.gguf3.0518.5212.2680.0019190.16174897.288%
Kai-3B-Instruct-i1-Q8_0.gguf3.0518.5212.2680.0019190.16174897.288%
Kai-3B-Instruct-Q6_K.gguf2.3576.5812.30550.0094040.36664994.435%
Kai-3B-Instruct-i1-Q6_K.gguf2.3576.5812.34860.0088420.52869994.605%
Kai-3B-Instruct-Q5_1.gguf2.1736.0712.46070.0225461.6205892.336%
Kai-3B-Instruct-i1-Q5_1.gguf2.1736.0712.39130.0155550.88786193.164%
Kai-3B-Instruct-Q5KM.gguf2.0625.7612.39320.0159532.0668493.315%
Kai-3B-Instruct-i1-Q5KM.gguf2.0625.7612.39740.0147121.2105493.344%
Kai-3B-Instruct-i1-Q5_0.gguf2.0145.6312.38450.0185821.781192.676%
Kai-3B-Instruct-Q5KS.gguf2.0095.6112.47050.0211122.2518892.477%
Kai-3B-Instruct-i1-Q5KS.gguf2.0095.6112.4220.0160981.0274293.198%
Kai-3B-Instruct-Q5_0.gguf2.0095.6112.53540.0245492.6475791.846%
Kai-3B-Instruct-i1-Q4_1.gguf1.8455.1612.66930.0392822.1726990.104%
Kai-3B-Instruct-Q4_1.gguf1.8455.1612.84110.0708939.7596387.274%
Kai-3B-Instruct-i1-Q4KM.gguf1.7844.9812.5620.0337912.3792990.693%
Kai-3B-Instruct-Q4KM.gguf1.7844.9812.55510.0393298.0895190.011%
Kai-3B-Instruct-IQ4_NL.gguf1.6974.7412.63490.047463.7583789.164%
Kai-3B-Instruct-Q4KS.gguf1.6934.7312.68810.0503177.1542188.889%
Kai-3B-Instruct-i1-Q4KS.gguf1.6934.7312.6720.0389762.3506290.141%
Kai-3B-Instruct-i1-Q4_0.gguf1.6874.7112.93180.0569144.9094288.242%
Kai-3B-Instruct-i1-IQ4_NL.gguf1.6864.7112.70290.0410412.8281489.995%
Kai-3B-Instruct-Q4_0.gguf1.6824.713.18310.0793595.3081386.546%
Kai-3B-Instruct-IQ4_XS.gguf1.6194.5212.66420.0485273.1169389.010%
Kai-3B-Instruct-i1-IQ4_XS.gguf1.6054.4812.73510.0421192.8166189.976%
Kai-3B-Instruct-Q3KL.gguf1.5744.413.22290.0953558.6383585.518%
Kai-3B-Instruct-i1-Q3KL.gguf1.5744.413.24770.0846685.7114386.163%
Kai-3B-Instruct-Q3KM.gguf1.4634.0913.34550.1126699.1984284.135%
Kai-3B-Instruct-i1-Q3KM.gguf1.4634.0913.40950.0959397.9367785.368%
Kai-3B-Instruct-i1-IQ3_M.gguf1.3683.8213.14810.1124376.4579984.307%"Note: Due to the ADS distillation method, this model is highly sensitive to quantization noise. Q80 or Q6K are strongly recommended for preserving both logical integrity and conversational alignment. Q4 variants may exhibit template collapse."
Kai-3B-Instruct-IQ3_M.gguf1.3683.8214.56930.2467137.2978177.711%
Kai-3B-Instruct-IQ3_S.gguf1.3393.7420.28510.62355714.944466.169%
Kai-3B-Instruct-i1-IQ3_S.gguf1.3393.7413.28230.1209756.1245183.724%
Kai-3B-Instruct-i1-Q3KS.gguf1.3343.7314.42790.19639611.924979.536%
Kai-3B-Instruct-Q3KS.gguf1.3343.7314.57530.2094710.276279.235%
Kai-3B-Instruct-i1-IQ3_XS.gguf1.2773.5713.57130.1498385.1909181.978%
Kai-3B-Instruct-i1-IQ3_XXS.gguf1.1813.314.49680.2183337.4113278.317%
Kai-3B-Instruct-i1-Q2_K.gguf1.1673.2617.05150.36285913.705473.511%
Kai-3B-Instruct-Q2_K.gguf1.1673.2618.4210.47169910.995570.276%
Kai-3B-Instruct-i1-Q2KS.gguf1.0963.0619.02030.471059.3998170.322%
Kai-3B-Instruct-i1-IQ2_M.gguf1.0482.9316.81790.3779148.0604872.505%
Kai-3B-Instruct-i1-IQ2_S.gguf0.9742.7218.96570.50757110.14668.855%
Kai-3B-Instruct-i1-IQ2_XS.gguf0.9462.6420.74340.6026312.284866.248%
Kai-3B-Instruct-i1-IQ2_XXS.gguf0.8682.4228.07160.91277220.855159.005%
Kai-3B-Instruct-i1-IQ1_M.gguf0.7762.1756.09381.7179716.768646.262%
Kai-3B-Instruct-i1-IQ1_S.gguf0.722.01142.1192.7124423.194935.970%

Kai-3B-Instruct

A 3B-parameter instruction-tuned language model optimized for reasoning, math, and code generation tasks, powered by our new ADS (Adaptive Dual-Search Distillation) technique.

Model Details

ModelKai-3B-Instruct
ArchitectureSmolLM3ForCausalLM
Parameters3B
Hidden size2048
Intermediate size11008
Layers36
Attention heads16 (4 KV heads, GQA)
Context length65536
Precisionbfloat16
Vocab size128,256

What is ADS?

Adaptive Dual-Search Distillation treats model fine-tuning as a constrained optimization problem inspired by Operations Research. The core mechanism is a dynamic loss function with a stateful dual penalty factor that adapts based on embedding space entropy — forcing the model to converge to high-confidence predictions at difficult reasoning points, without modifying the model architecture.

Benchmark Results

[image]

General (5-shot, log-likelihood)

ModelParamsMMLUARC-c (acc_norm)HellaSwag (acc_norm)PIQA (acc_norm)
TinyLlama1.1B~26.0%~33.0%~60.0%~71.0%
SmolLM21.7B~35.0%~38.0%~65.0%~74.0%
Llama-2-7B7B45.3%46.2%77.2%79.8%
Gemma-2-2B2.6B~52.0%~53.0%75.0%~78.0%
Kai-3B-Instruct3B53.62%51.88%69.53%77.53%
Qwen2.5-3B3B~63.0%~55.0%~73.0%~80.0%

Code Generation — HumanEval (Pass@1, 0-shot)

ModelParamsHumanEval (Pass@1)Notes
Llama-2-7B7B~12.8%3x overtake — smaller model, far better code
SmolLM2-1.7B1.7B~25.0%ADS delivers +14pp pure gain
Gemma-2-2B2B~30.0%Surpasses Google's heavily distilled 2B flagship
Kai-3B-Instruct3B39.02%ADS topological pruning, full pipeline
GPT-3.5 (Legacy)175B~48.0%Kai-3B trails the original GPT-3.5 by only ~9pp

Math — GSM8K (0-shot)

ModelParamsGSM8K (exact_match)
Kai-3B-Instruct3B39.27%

Key Observations

  1. 1.Surpasses Llama-2-7B: Kai-3B outperforms Llama-2-7B on MMLU (+8.3pp) and ARC-Challenge (+5.7pp) with less than half the parameters — a 7B model decisively beaten by a 3B distilled model.
  1. 1.Competitive with Gemma-2-2B: Matches or exceeds Google's Gemma-2-2B on MMLU (+1.6pp) and PIQA, despite Gemma being trained with significantly more compute.
  1. 1.HellaSwag: At 69.53%, Kai-3B surpasses all sub-2B models by a wide margin and trails the compute-heavy Qwen2.5-3B by only ~3.5pp.
  1. 1.PIQA: At 77.53%, Kai-3B nearly matches Gemma-2-2B (~78.0%) and approaches the 3B-class ceiling set by Qwen2.5-3B (~80.0%).

Usage

python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model = AutoModelForCausalLM.from_pretrained(
    "NoesisLab/Kai-3B-Instruct",
    torch_dtype=torch.bfloat16,
)
tokenizer = AutoTokenizer.from_pretrained("NoesisLab/Kai-3B-Instruct")

messages = [{"role": "user", "content": "What is 25 * 4?"}]
input_ids = tokenizer.apply_chat_template(messages, return_tensors="pt")
output = model.generate(input_ids, max_new_tokens=256)
print(tokenizer.decode(output[0], skip_special_tokens=True))

Citation

bibtex
@misc{noesislab2026kai3b,
  title={Kai-3B-Instruct},
  author={NoesisLab},
  year={2026},
  url={https://huggingface.co/NoesisLab/Kai-3B-Instruct}
}

License

Apache 2.0