Nanite-Labs/nanites-smoky-1b-chat
nanites-smoky-1b-chat
A Llama-3.2-1B-Instruct model fine-tuned to play the role of an old resident of the Great Smoky Mountains of the 1930s in conversation. You talk to it; the old-timer answers in period voice with regional vocabulary.
This is the chat-tuned sibling of the CosyVoice2 gender-pooled speech voices (Earl / male at Nanite-Labs/nanites-smoky-speech-earl, Ethel / female at Nanite-Labs/nanites-smoky-speech-ethel). The chat model provides the language behind the TTS; pair them for a text-and-speech experience.
What it does
Given a modern-English question about life in the Smokies (work, family, food, music, religion, hardships, the mountains themselves), the model answers in 1-3 sentences in the voice of an old Smoky Mountain resident. The voice uses regional vocabulary ("warn't", "pore", "y'all", "holler", "crick", "taters", "a-going", "nigh onto", etc.) and avoids modern references.
This is a small (1B) finetune on a small (~93 Q&A in 84 train / 9 val) curated dataset, not a comprehensive knowledge model. It plays the character convincingly for the regional register; it does not answer factual questions about the Smokies in general.
What it doesn't do
- It is not an authoritative history source. The training data is interviews from Joseph Sargent Hall's 1939 field recordings; the model's "facts" reflect those speakers' lived experience, not academic history.
- It does not represent any single Hall speaker — the dataset was distilled by a teacher LLM that paraphrased the recordings in a generalized regional voice. No voice clone, no single persona.
- It cannot generate audio. For speech, use the CosyVoice2 voices (
experts/smoky/cosyvoice_infer.py --sft_spk_id ...).
Training
- Base: unsloth/llama-3.2-1b-instruct-unsloth-bnb-4bit (the 4-bit quantized Llama-3.2-1B-Instruct; the adapter fine-tunes on top)
- Method: QLoRA via unsloth (4-bit base, LoRA r=16 on all attention + MLP projections)
- Data: 84 train / 9 val Alpaca-format Q&A pairs (
experts/smoky/data/splits_chat/). The seed is a small hand-authored set; the bulk is teacher-distilled from the Hall transcripts via a local ollama-served LLM (nemotron-3-nano:30b). The teacher is asked to paraphrase in a generalized regional voice, so the chat model does not reproduce the speakers' specific words or voices. - Context: 1024 tokens
- Output:
models/smoky/run_chat/adapter/(LoRA adapter)
How to use
The adapter is published here. To run locally:
import torch
from unsloth import FastLanguageModel
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="Nanite-Labs/nanites-smoky-1b-chat",
max_seq_length=1024,
dtype=None,
load_in_4bit=True,
)
FastLanguageModel.for_inference(model)
prompt = """Below is an instruction that describes a task. Write a response that appropriately completes the request.
### Instruction:
You are an old resident of the Great Smoky Mountains of the 1930s. A visitor has just asked you a question. Answer in the voice of a Smoky Mountain old-timer — 1-3 sentences, regional vocabulary, period-appropriate. Stay in character.
### Input:
How was life growing up around here?
### Response:
"""
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
out = model.generate(
**inputs,
max_new_tokens=200,
do_sample=True,
temperature=0.7,
pad_token_id=tokenizer.eos_token_id,
)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))Or run the included REPL:
python -m experts.smoky.chatPairing with TTS
For a speech-enabled experience, synthesize the chat model's output with the CosyVoice2 gender-pooled voices:
python experts/smoky/cosyvoice_infer.py \
--model_dir experts/smoky/data/cosyvoice/voices/smoky-speech-male \
--text "$(python -m experts.smoky.chat --noninteractive 'What was sharecropping like?')" \
--sft_spk_id "<|male_pool|>" \
--out out.wavProvenance
- Base model: meta-llama/Llama-3.2-1B-Instruct via unsloth's 4-bit quantization
- Training data: Joseph Sargent Hall Collection, University of South Carolina Southern Appalachian English archive (https://appalachian-english.library.sc.edu/transcripts.html). The audio recordings (1939) are public domain. The transcripts are copyrighted to Michael Montgomery and Paul Reed (2017) and were used locally only — they are not redistributed in this repo. The Q&A training pairs were generated from the transcripts by a teacher LLM and are also not the raw transcript text.
License
Code & recipe: Apache 2.0 Base model: see Meta Llama 3.2 Community License Training data (audio): public domain (1939) Training data (transcripts): not redistributed Training data (Q&A): generated by local teacher LLM, not raw transcript
