CoolFace
Modelpublic

NANI-Nithin/north-mini-code-gguf

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
1likes116downloads
Model Card

North-Mini-Code-1.0-GGUF

GGUF conversions and quantizations of CohereLabs/North-Mini-Code-1.0 for use with:

  • —llama.cpp
  • —LM Studio
  • —Ollama
  • —Jan
  • —KoboldCpp
  • —Text Generation WebUI
  • —Open WebUI
  • —Other GGUF-compatible runtimes

About the Model

North-Mini-Code-1.0 is a code-focused Mixture-of-Experts (MoE) model released by CohereLabs.

This repository provides ready-to-use GGUF conversions for local inference across a range of hardware configurations.


Available Files

Full Precision

  • —North-Mini-Code-1.0-F16.gguf

Quantized Versions

  • —North-Mini-Code-1.0-Q4_K_M.gguf
  • —North-Mini-Code-1.0-Q5_K_M.gguf
  • —North-Mini-Code-1.0-Q6_K.gguf
  • —North-Mini-Code-1.0-Q8_0.gguf

Recommended Quantization

For most users:

text
North-Mini-Code-1.0-Q4_K_M.gguf

It offers the best balance of:

  • —Quality
  • —Memory usage
  • —Inference speed

If you have more available RAM/VRAM, consider:

text
North-Mini-Code-1.0-Q5_K_M.gguf

or

text
North-Mini-Code-1.0-Q6_K.gguf

for slightly higher output quality.


Approximate File Sizes

text
F16      ~60+ GB
Q4_K_M   ~20 GB
Q5_K_M   ~23 GB
Q6_K     ~27 GB
Q8_0     ~34 GB

Actual sizes may vary slightly depending on conversion tooling versions.


Usage

llama.cpp

Prompt mode:

bash
./llama-cli \
  -m North-Mini-Code-1.0-Q4_K_M.gguf \
  -p "Write a Python function that reverses a linked list."

Chat mode:

bash
./llama-cli \
  -m North-Mini-Code-1.0-Q4_K_M.gguf \
  -cnv

LM Studio

  1. 1.Download your preferred GGUF file.
  2. 2.Open LM Studio.
  3. 3.Import the model.
  4. 4.Start chatting.

Ollama

Create a Modelfile:

text
FROM North-Mini-Code-1.0-Q4_K_M.gguf

Create the model:

bash
ollama create north-mini-code -f Modelfile

Run it:

bash
ollama run north-mini-code

Hardware Recommendations

Q4KM

Recommended minimum:

text
24 GB RAM

Q5KM

Recommended minimum:

text
32 GB RAM

Q6_K

Recommended minimum:

text
32-40 GB RAM

Q8_0

Recommended minimum:

text
48+ GB RAM

F16

Recommended minimum:

text
80+ GB RAM

Prompting Tips

This model is optimized for programming-related tasks.

Example prompts:

text
Implement a fast Rust HTTP server.
text
Explain this C++ compiler error.
text
Write comprehensive unit tests for the following Python code.
text
Convert this JavaScript function to TypeScript.
text
Optimize this SQL query.

Base Model

Base model:

text
CohereLabs/North-Mini-Code-1.0

All training, architecture, benchmarks, licensing terms, and usage restrictions belong to the original model authors.

Please refer to the original repository for official documentation and licensing information.


Conversion Details

Converted using:

text
llama.cpp

Generated quantizations:

text
F16
Q4_K_M
Q5_K_M
Q6_K
Q8_0

A tokenizer compatibility workaround was applied during conversion to support current GGUF conversion tooling.


Credits

  • —Base Model: CohereLabs
  • —GGUF Conversion & Quantization: NANI-Nithin
  • —Tooling: llama.cpp

Repository

👉 https://huggingface.co/NANI-Nithin/north-mini-code-gguf