AnkitAI/Parable-SmolLM3-3B-Claude-Fable-5
Parable-SmolLM3-3B-Claude-Fable-5
Part of the Parable series: small local LLMs fine-tuned on genuine agent traces. This is HuggingFaceTB/SmolLM3-3B tuned on real Claude Fable 5 agent transcripts so its step-by-step reasoning voice carries into local use.
Quantized GGUF builds for llama.cpp / LM Studio / Ollama: Parable-SmolLM3-3B-Claude-Fable-5-GGUF
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "AnkitAI/Parable-SmolLM3-3B-Claude-Fable-5"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="auto", device_map="auto")
msgs = [{"role": "user", "content": "Write a python function that reverses a string."}]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=400, temperature=0.6)
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True))Output begins with a <think>...</think> reasoning block, then the answer. Parse and strip the think block before showing text to end users. The chat template identifies the model as "Parable, a coding assistant that reasons before it answers."
Model details
- Base: HuggingFaceTB/SmolLM3-3B (3B, Apache-2.0, 64k context)
- Method: MLX QLoRA on a 4-bit quantized base, 16 layers adapted, 6.7M trainable parameters (0.218%); this repo holds the dequantized F16 merge as safetensors
- Data: 4,076 training rows from real Claude Fable 5 agent-session traces plus gpt5.5-terminal transcripts, prepared at a 4,096-token window (268 over-length rows dropped; 226/226 rows held out for validation/test)
- Schedule: 1,200-iteration budget across a paused-and-resumed run; best checkpoint selected on validation loss (iteration 200 of the final segment, val 1.154)
Evaluation
The tuned model fits the Fable-5 reasoning distribution 41% better by held-out loss on a 226-row test split never seen in training. That is the honest headline for what this fine-tune does; we do not claim general benchmark gains.
This lane trains on trace data without a replay mix, so impact on general coding benchmarks is unmeasured here. The series' technical report (DOI: 10.5281/zenodo.21676407) documents why that matters and what replay does about it.
Limitations
- Training ran on a 4-bit quantized base (16 GB M1 constraint); the F16 merge cannot exceed 4-bit-base quality.
- Modest scale: one seed, loss-based evaluation, no external benchmark run for this model yet.
- Not trained for: multi-file repo navigation, vision, non-English.
- Inherits SmolLM3-3B's knowledge cutoff. Treat generated commands as drafts to review.
Provenance & licensing
Fine-tuned from HuggingFaceTB/SmolLM3-3B (Apache-2.0). Training data: Glint-Research/Fable-5-traces (AGPL-3.0) and Roman1111111/gpt5.5-terminal (MIT). Because those traces originate from third-party assistants, the providers' terms may apply to downstream training and distillation. If you plan to build on this model commercially, confirm your use aligns with those terms.
Support the Project
If this model is useful in your work, you can support independent research:
<p align="left"> <a href="https://www.buymeacoffee.com/AnkitAI" target="_blank"><img src="https://cdn.buymeacoffee.com/buttons/v2/default-yellow.png" alt="Buy Me a Coffee" height="60" width="217" /></a> </p>
Citation
@misc{aglawe2026parable,
author = {Aglawe, Ankit},
title = {Agent-Trace Fine-Tuning of Small Language Models under Constrained Compute},
year = {2026},
doi = {10.5281/zenodo.21676407},
url = {https://doi.org/10.5281/zenodo.21676407}
}Acknowledgements
The SmolLM3 team at Hugging Face for the base model; Glint-Research and Roman1111111 for the trace datasets; empero-ai for the recipe this series iterates on.
