AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5-GGUF
<picture> <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/ankit-aglawe/parable-assets/main/parableheaderdark.png"> <img alt="Parable" src="https://raw.githubusercontent.com/ankit-aglawe/parable-assets/main/parable_header.png"> </picture>
๐ชถ Parable-Granite-3B v2 โ trained on genuine Claude Fable 5 agent traces
A tiny local model that thinks before it answers โ planning, reasoning, and terminal instincts distilled from real agent sessions.
~3 GB of RAM is all you need. Laptop, old GPU, Raspberry-Pi-class boxes with swap โ the Q4 build runs anywhere. One command and you have a private, offline reasoning model on your machine: ``bash ollama run hf.co/AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5-GGUF:Q4_K_M ``The headline โ v2 is a different model
v2 is a full retrain: 13ร more genuine Fable 5 trace data (11,574 sessions, 16.8M tokens โ corpus published) and a rebuilt recipe (completion-masked loss, replay mixing, benchmark-gated checkpoints, seed-averaged weights).
Clean answers, structured reasoning, agent instincts โ and the transcript artifacts that leaked into v1's replies are gone. One trade, made on purpose: raw HumanEval-style function synthesis stays the base model's turf (81.7 vs 70.1) โ v2 spends that capacity on agent behavior instead, and spends half as much as v1 did. Measurement notes below. ๐
Announcements
๐ Same links, new model. v2 replaces v1 in place โ every existing Ollama command, script, and bookmark now serves v2. No migration, nothing to change.
๐ฎ v3 is already training. Rejection-sampled SFT: thousands of candidate solutions generated against executable tests, only verified passers enter the corpus. The goal is simple โ above-base agent capability, not just clean behavior. Follow AnkitAI for the drop.
๐ฆ Full family. This 3B is the smallest Parable. Need more headroom? 8B Granite, 8B Qwen, 4B Qwen โ same recipe, no matter your hardware.
Pick your size
Intelligence per gigabyte: the Q4KM build scores 70.1 HumanEval in 2.1 GB โ ~33 pts/GB; an 8B-class Q4 needs ~5 GB for its score. If RAM is your constraint, this is the family's density sweet spot.
Full-precision safetensors (vLLM, transformers, further fine-tuning): Parable-Granite-4.1-3B-Claude-Fable-5
How to run it
Ollama (chat template ships inside the GGUF โ zero config):
ollama run parable/granite4.1-fable:3b
# or straight from this repo:
ollama run hf.co/AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5-GGUF:Q4_K_Mllama.cpp:
llama-cli -m Parable-Granite-4.1-3B-Claude-Fable-5-GGUF-Q4_K_M.gguf --jinja \
-p "Write a bash one-liner to find the 10 largest files in a directory tree."LM Studio / Jan / Open WebUI: search "parable" in-app, or paste this repo URL.
Python (llama-cpp-python):
from llama_cpp import Llama
llm = Llama.from_pretrained(
repo_id="AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5-GGUF",
filename="*Q4_K_M.gguf", n_ctx=8192,
)
out = llm.create_chat_completion(
messages=[{"role": "user", "content": "Write a Python function that retries an HTTP request with exponential backoff."}],
max_tokens=3000, temperature=0.7,
)
print(out["choices"][0]["message"]["content"])Thinking mode
Every answer opens with a <think>...</think> reasoning block โ that's the Fable 5 heritage. llama.cpp's --jinja mode separates it automatically; strip it before showing replies to end users. Sampling: temperature 0.7, topp 0.95, and budget `maxtokens` generously (2500+) โ trace-trained models think at length before answering.
Measurement notes
All numbers: identical llama.cpp harness, greedy decoding, Q4KM, base model measured on the same instrument. We train multiple seeds and ship the weight-average โ single-run scores at 3B swing ยฑ3 points on GPU nondeterminism alone, so most cards report their luckiest run; we ship the average and report the shipped weights' own numbers. Raw eval outputs live in this repo.
Which model should you use? Pure single-function code completion โ the base model is genuinely strong there. Explanations, debugging, terminal workflows, structured reasoning, agent-style tasks โ that's what Parable is trained on, and where v2 shines.
๐ Same prompt, side by side
Real outputs, both models at Q4KM, temperature 0.7 โ unedited except length.
Prompt: "Make this more idiomatic:" result = []; for x in items: if x.active == True: result.append(x.name.upper())
Prompt: "My Python script fails with 'RecursionError: maximum recursion depth exceeded' in a JSON parser I wrote. What are the likely causes and the standard fix?"
The pattern from real agent traces: answer first, structure over prose, no padding. (Where the base is stronger โ raw single-function synthesis โ is stated plainly in the measurement notes above.)
What's new in v2 (training)
The recipe follows our ongoing tech report (in preparation):
- Completion-only loss masking (Hermes 3, Tรผlu 3) โ loss on assistant tokens only, so the model learns to answer, not to imitate transcripts
- 30% replay mix of general instruction data (Luo et al., Biderman et al.) โ the anti-forgetting lever
- Session re-segmentation + sanitization โ why v1 sometimes leaked agent JSON into normal chat, and v2 never does (0/34)
- Benchmark-gated checkpoints (Dong et al.) instead of fixed epochs
- Seed-averaged weights (model soups, Wortsman et al.) โ we ship the average of multiple runs, not the lottery winner
With Claude Fable 5 now retired, genuine self-authored Fable traces are a fixed, non-renewable corpus. Unlike most models in this niche, our full training corpus is public: AnkitAI/parable-corpus-v2 โ deduplicated, quality-gated, provenance-tagged.
Good to know
- Fine-tuned at 2,048-token sequences; the base 128K context stays available, fine-tuned behavior is strongest in the opening turns.
- Not trained for: multi-file repo navigation, vision, non-English.
- Inherits Granite-4.1-3B's knowledge cutoff. Treat generated commands as drafts to review.
Evaluation
Function calling (BFCL V3, AST subset)
Measured 2026-07-29: bfcl-eval at gorilla main, prompting mode, Q4KM GGUFs served by llama.cpp on a T4, base and Parable under the identical harness. Categories: simplepython / multiple / parallel / parallelmultiple (400/200/200/200 items). Raw generations and score files: parable-v2-artifacts under verify/bfcl/.
For tool-calling workloads, use the base model; this variant is built for reasoning prose. The drop has a specific mechanism: sampled generations show the model intermittently answering with args-only tool-call JSON (for example {"base": 10, "height": 5}) instead of a function call, which the AST scorer rejects. That is trace-scaffolding format bleeding into standalone tasks, the failure mode the series paper names session leakage (Section 6 of the report). The reasoning-voice strengths this variant trains for are unaffected on prose tasks.
Base & license
Weights: Apache-2.0 (inherited from ibm-granite/granite-4.1-3b). Training data: Fable-5-traces AGPL-3.0, gpt5.5-terminal MIT โ since traces originate from third-party assistants, their terms may apply to downstream training; check before commercial distillation.
Get Parable
Citation
The recipe, evaluation methodology and failure analysis behind this model are documented in the tech report:
Aglawe, A. (2026). Agent-Trace Fine-Tuning of Small Language Models under Constrained Compute. Zenodo. doi:10.5281/zenodo.21676407
@misc{aglawe2026agenttrace,
author = {Aglawe, Ankit},
title = {Agent-Trace Fine-Tuning of Small Language Models under Constrained Compute},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.21676407},
url = {https://doi.org/10.5281/zenodo.21676407}
}Acknowledgements
Glint-Research & Roman1111111 for the open trace data ยท IBM Granite for the base ยท empero-ai whose Qwable recipe inspired the series ยท llama.cpp
Version history
- v2 (2026-07-16) โ this release. 13ร corpus, rebuilt recipe, seed-averaged weights, zero leakage.
- v1 (2026-07) โ initial release, 857-row corpus. Preserved as repo revision history.
Three gigabytes. Real Fable 5 reasoning. Yours, offline, right now.
ollama run hf.co/AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5-GGUF:Q4_K_MMore on the Parable models: ankitaglawe.com/parable
