CoolFace
Modelpublic

glogwa68/granite-4.0-h-1b-DISTILL-glm-4.7-think

sourceHugging Faceapache-2.0updated 9mo agoView on Hugging Face
1likes64downloads
Model Card

granite-4.0-h-1b-DISTILL-glm-4.7-think

This model is a fine-tuned version of ibm-granite/granite-4.0-h-1b trained on conversational data.

Model Details

  • —Base Model: ibm-granite/granite-4.0-h-1b
  • —Fine-tuning Dataset: TeichAI/glm-4.7-2000x
  • —Training Loss: 0.6364
  • —Context Length: 1048576 tokens

Quantized Versions (GGUF)

🔗 GGUF versions available here: [granite-4.0-h-1b-DISTILL-glm-4.7-think-GGUF](https://huggingface.co/glogwa68/granite-4.0-h-1b-DISTILL-glm-4.7-think-GGUF)

FormatSizeUse Case
Q2_KSmallestLow memory, reduced quality
Q4KMRecommendedBest balance
Q5KMGoodHigher quality
Q8_0LargeNear lossless
F16LargestOriginal precision

Usage

Transformers

python
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("glogwa68/granite-4.0-h-1b-DISTILL-glm-4.7-think")
tokenizer = AutoTokenizer.from_pretrained("glogwa68/granite-4.0-h-1b-DISTILL-glm-4.7-think")

messages = [{"role": "user", "content": "Hello, how are you?"}]
inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True)
outputs = model.generate(inputs, max_new_tokens=256)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Ollama (GGUF)

bash
ollama run hf.co/glogwa68/granite-4.0-h-1b-DISTILL-glm-4.7-think-GGUF:Q4_K_M

llama.cpp

bash
llama-cli --hf-repo glogwa68/granite-4.0-h-1b-DISTILL-glm-4.7-think-GGUF --hf-file granite-4.0-h-1b-distill-glm-4.7-think-q4_k_m.gguf -p "Hello"

Training Details

  • —Epochs: 3
  • —Learning Rate: 2e-5
  • —Batch Size: 1 (with gradient accumulation)
  • —Precision: FP16
  • —Hardware: Multi-GPU with DeepSpeed ZeRO-3

License

Apache 2.0