CoolFace
Modelpublic

tensorblock/ServiceNow-AI_Apriel-Nemotron-15b-Thinker-GGUF

sourceHugging Facemitupdated 8mo agoView on Hugging Face
1likes156downloads
Model Card

<div style="width: auto; margin-left: auto; margin-right: auto"> <img src="https://i.imgur.com/jC7kdl8.jpeg" alt="TensorBlock" style="width: 100%; min-width: 400px; display: block; margin: auto;"> </div>

![Website](https://tensorblock.co) ![Twitter](https://twitter.com/tensorblock_aoi) ![Discord](https://discord.gg/Ej5NmeHFf2) ![GitHub](https://github.com/TensorBlock) ![Telegram](https://t.me/TensorBlock)

ServiceNow-AI/Apriel-Nemotron-15b-Thinker - GGUF

<div style="text-align: left; margin: 20px 0;"> <a href="https://discord.com/invite/Ej5NmeHFf2" style="display: inline-block; padding: 10px 20px; background-color: #5865F2; color: white; text-decoration: none; border-radius: 5px; font-weight: bold;"> Join our Discord to learn more about what we're building โ†— </a> </div>

This repo contains GGUF format model files for ServiceNow-AI/Apriel-Nemotron-15b-Thinker.

The files were quantized using machines provided by TensorBlock, and they are compatible with llama.cpp as of commit b5753.

Our projects

<table border="1" cellspacing="0" cellpadding="10"> <tr> <th colspan="2" style="font-size: 25px;">Forge</th> </tr> <tr> <th colspan="2"> <img src="https://imgur.com/faI5UKh.jpeg" alt="Forge Project" width="900"/> </th> </tr> <tr> <th colspan="2">An OpenAI-compatible multi-provider routing layer.</th> </tr> <tr> <th colspan="2"> <a href="https://github.com/TensorBlock/forge" target="_blank" style=" display: inline-block; padding: 8px 16px; background-color: #FF7F50; color: white; text-decoration: none; border-radius: 6px; font-weight: bold; font-family: sans-serif; ">๐Ÿš€ Try it now! ๐Ÿš€</a> </th> </tr>

<tr> <th style="font-size: 25px;">Awesome MCP Servers</th> <th style="font-size: 25px;">TensorBlock Studio</th> </tr> <tr> <th><img src="https://imgur.com/2Xov7B7.jpeg" alt="MCP Servers" width="450"/></th> <th><img src="https://imgur.com/pJcmF5u.jpeg" alt="Studio" width="450"/></th> </tr> <tr> <th>A comprehensive collection of Model Context Protocol (MCP) servers.</th> <th>A lightweight, open, and extensible multi-LLM interaction studio.</th> </tr> <tr> <th> <a href="https://github.com/TensorBlock/awesome-mcp-servers" target="blank" style=" display: inline-block; padding: 8px 16px; background-color: #FF7F50; color: white; text-decoration: none; border-radius: 6px; font-weight: bold; font-family: sans-serif; ">๐Ÿ‘€ See what we built ๐Ÿ‘€</a> </th> <th> <a href="https://github.com/TensorBlock/TensorBlock-Studio" target="blank" style=" display: inline-block; padding: 8px 16px; background-color: #FF7F50; color: white; text-decoration: none; border-radius: 6px; font-weight: bold; font-family: sans-serif; ">๐Ÿ‘€ See what we built ๐Ÿ‘€</a> </th> </tr> </table>

Prompt template

<|system|>
You are a thoughtful and systematic AI assistant built by ServiceNow Language Models (SLAM) lab. Before providing an answer, analyze the problem carefully and present your reasoning step by step. After explaining your thought process, provide the final solution in the following format: [BEGIN FINAL RESPONSE] ... [END FINAL RESPONSE].

{system_prompt}
<|end|>
<|user|>
{prompt}
<|end|>
<|assistant|>
Here are my reasoning steps:

Model file specification

FilenameQuant typeFile SizeDescription
Apriel-Nemotron-15b-Thinker-Q2_K.ggufQ2_K5.794 GBsmallest, significant quality loss - not recommended for most purposes
Apriel-Nemotron-15b-Thinker-Q3_K_S.ggufQ3KS6.706 GBvery small, high quality loss
Apriel-Nemotron-15b-Thinker-Q3_K_M.ggufQ3KM7.396 GBvery small, high quality loss
Apriel-Nemotron-15b-Thinker-Q3_K_L.ggufQ3KL7.990 GBsmall, substantial quality loss
Apriel-Nemotron-15b-Thinker-Q4_0.ggufQ4_08.606 GBlegacy; small, very high quality loss - prefer using Q3KM
Apriel-Nemotron-15b-Thinker-Q4_K_S.ggufQ4KS8.663 GBsmall, greater quality loss
Apriel-Nemotron-15b-Thinker-Q4_K_M.ggufQ4KM9.113 GBmedium, balanced quality - recommended
Apriel-Nemotron-15b-Thinker-Q5_0.ggufQ5_010.393 GBlegacy; medium, balanced quality - prefer using Q4KM
Apriel-Nemotron-15b-Thinker-Q5_K_S.ggufQ5KS10.393 GBlarge, low quality loss - recommended
Apriel-Nemotron-15b-Thinker-Q5_K_M.ggufQ5KM10.655 GBlarge, very low quality loss - recommended
Apriel-Nemotron-15b-Thinker-Q6_K.ggufQ6_K12.293 GBvery large, extremely low quality loss
Apriel-Nemotron-15b-Thinker-Q8_0.ggufQ8_015.919 GBvery large, extremely low quality loss - not recommended

Downloading instruction

Command line

Firstly, install Huggingface Client

shell
pip install -U "huggingface_hub[cli]"

Then, downoad the individual model file the a local directory

shell
huggingface-cli download tensorblock/ServiceNow-AI_Apriel-Nemotron-15b-Thinker-GGUF --include "Apriel-Nemotron-15b-Thinker-Q2_K.gguf" --local-dir MY_LOCAL_DIR

If you wanna download multiple model files with a pattern (e.g., *Q4_K*gguf), you can try:

shell
huggingface-cli download tensorblock/ServiceNow-AI_Apriel-Nemotron-15b-Thinker-GGUF --local-dir MY_LOCAL_DIR --local-dir-use-symlinks False --include='*Q4_K*gguf'