CoolFace
Modelpublic

tensorblock/THU-KEG_LongWriter-Zero-32B-GGUF

sourceHugging Faceapache-2.0updated 8mo agoView on Hugging Face
1likes114downloads
Model Card

<div style="width: auto; margin-left: auto; margin-right: auto"> <img src="https://i.imgur.com/jC7kdl8.jpeg" alt="TensorBlock" style="width: 100%; min-width: 400px; display: block; margin: auto;"> </div>

![Website](https://tensorblock.co) ![Twitter](https://twitter.com/tensorblock_aoi) ![Discord](https://discord.gg/Ej5NmeHFf2) ![GitHub](https://github.com/TensorBlock) ![Telegram](https://t.me/TensorBlock)

THU-KEG/LongWriter-Zero-32B - GGUF

<div style="text-align: left; margin: 20px 0;"> <a href="https://discord.com/invite/Ej5NmeHFf2" style="display: inline-block; padding: 10px 20px; background-color: #5865F2; color: white; text-decoration: none; border-radius: 5px; font-weight: bold;"> Join our Discord to learn more about what we're building โ†— </a> </div>

This repo contains GGUF format model files for THU-KEG/LongWriter-Zero-32B.

The files were quantized using machines provided by TensorBlock, and they are compatible with llama.cpp as of commit b5753.

Our projects

<table border="1" cellspacing="0" cellpadding="10"> <tr> <th colspan="2" style="font-size: 25px;">Forge</th> </tr> <tr> <th colspan="2"> <img src="https://imgur.com/faI5UKh.jpeg" alt="Forge Project" width="900"/> </th> </tr> <tr> <th colspan="2">An OpenAI-compatible multi-provider routing layer.</th> </tr> <tr> <th colspan="2"> <a href="https://github.com/TensorBlock/forge" target="_blank" style=" display: inline-block; padding: 8px 16px; background-color: #FF7F50; color: white; text-decoration: none; border-radius: 6px; font-weight: bold; font-family: sans-serif; ">๐Ÿš€ Try it now! ๐Ÿš€</a> </th> </tr>

<tr> <th style="font-size: 25px;">Awesome MCP Servers</th> <th style="font-size: 25px;">TensorBlock Studio</th> </tr> <tr> <th><img src="https://imgur.com/2Xov7B7.jpeg" alt="MCP Servers" width="450"/></th> <th><img src="https://imgur.com/pJcmF5u.jpeg" alt="Studio" width="450"/></th> </tr> <tr> <th>A comprehensive collection of Model Context Protocol (MCP) servers.</th> <th>A lightweight, open, and extensible multi-LLM interaction studio.</th> </tr> <tr> <th> <a href="https://github.com/TensorBlock/awesome-mcp-servers" target="blank" style=" display: inline-block; padding: 8px 16px; background-color: #FF7F50; color: white; text-decoration: none; border-radius: 6px; font-weight: bold; font-family: sans-serif; ">๐Ÿ‘€ See what we built ๐Ÿ‘€</a> </th> <th> <a href="https://github.com/TensorBlock/TensorBlock-Studio" target="blank" style=" display: inline-block; padding: 8px 16px; background-color: #FF7F50; color: white; text-decoration: none; border-radius: 6px; font-weight: bold; font-family: sans-serif; ">๐Ÿ‘€ See what we built ๐Ÿ‘€</a> </th> </tr> </table>

Prompt template

A conversation between the user and the assistant. The user provides a writing/general task, and the assistant completes it. The assistant first deeply thinks through the writing/answering process in their mind before providing the final written work to the user. The assistant should engage in comprehensive and in-depth planning to ensure that every aspect of the writing/general task is detailed and well-structured. If there is any uncertainty or ambiguity in the writing request, the assistant should reflect, ask themselves clarifying questions, and explore multiple writing approaches to ensure the final output meets the highest quality standards. Since writing is both a creative and structured task, the assistant should analyze it from multiple perspectives, considering coherence, clarity, style, tone, audience, purpose, etc.. Additionally, the assistant should review and refine the work to enhance its expressiveness. The writing thought process and the final written work should be enclosed within <think> </think> and <answer> </answer> tags, respectively, as shown below: <think>A comprehensive strategy for writing that encompasses detailed planning and structural designโ€”including brainstorming, outlining, style selection, audience adaptation, self-reflection, quality assurance, etc..</think> <answer>The final written work after thorough optimization and refinement.</answer> <|user|>: {system_prompt} <|assistant|>:

Model file specification

FilenameQuant typeFile SizeDescription
LongWriter-Zero-32B-Q2_K.ggufQ2_K12.313 GBsmallest, significant quality loss - not recommended for most purposes
LongWriter-Zero-32B-Q3_K_S.ggufQ3KS14.392 GBvery small, high quality loss
LongWriter-Zero-32B-Q3_K_M.ggufQ3KM15.935 GBvery small, high quality loss
LongWriter-Zero-32B-Q3_K_L.ggufQ3KL17.247 GBsmall, substantial quality loss
LongWriter-Zero-32B-Q4_0.ggufQ4_018.640 GBlegacy; small, very high quality loss - prefer using Q3KM
LongWriter-Zero-32B-Q4_K_S.ggufQ4KS18.784 GBsmall, greater quality loss
LongWriter-Zero-32B-Q4_K_M.ggufQ4KM19.851 GBmedium, balanced quality - recommended
LongWriter-Zero-32B-Q5_0.ggufQ5_022.638 GBlegacy; medium, balanced quality - prefer using Q4KM
LongWriter-Zero-32B-Q5_K_S.ggufQ5KS22.638 GBlarge, low quality loss - recommended
LongWriter-Zero-32B-Q5_K_M.ggufQ5KM23.262 GBlarge, very low quality loss - recommended
LongWriter-Zero-32B-Q6_K.ggufQ6_K26.886 GBvery large, extremely low quality loss
LongWriter-Zero-32B-Q8_0.ggufQ8_034.821 GBvery large, extremely low quality loss - not recommended

Downloading instruction

Command line

Firstly, install Huggingface Client

shell
pip install -U "huggingface_hub[cli]"

Then, downoad the individual model file the a local directory

shell
huggingface-cli download tensorblock/THU-KEG_LongWriter-Zero-32B-GGUF --include "LongWriter-Zero-32B-Q2_K.gguf" --local-dir MY_LOCAL_DIR

If you wanna download multiple model files with a pattern (e.g., *Q4_K*gguf), you can try:

shell
huggingface-cli download tensorblock/THU-KEG_LongWriter-Zero-32B-GGUF --local-dir MY_LOCAL_DIR --local-dir-use-symlinks False --include='*Q4_K*gguf'