VertexAGI/amethyst-1-mini
Amethyst 1 Mini
Amethyst 1 Mini is a general-purpose chat and instruction-following model, fine-tuned from Gemma 3 4B IT using LoRA on a small, high-quality distilled instruction dataset. It is the first model in the Amethyst family — an early, small-scale release focused on validating an end-to-end distillation → fine-tune → evaluate pipeline on consumer hardware.
Model Details
Training Data
Amethyst 1 Mini was fine-tuned on 1,122 instruction/response pairs (1,082 train / 40 validation), synthetically generated via knowledge distillation from `nvidia/nemotron-3-super-120b-a12b` (Nemotron-3-Super, a 120B-parameter MoE model, ~12B active) through the OpenRouter API.
The dataset spans a deliberately broad set of general-chat categories to encourage well-rounded conversational ability rather than narrow task performance:
- Explanation & misconceptions
- Reasoning & math reasoning
- Code generation
- Extraction & structured output
- Planning
- Roleplay & creative writing
- Translation
- Sentiment classification
- Brainstorming
Each example was generated from a unique prompt, distilled from the teacher model, then filtered/cleaned before training.
Training Procedure
- Method: Supervised fine-tuning via LoRA (rank 8, dropout 0.0, scale 20.0)
- Optimizer: Adam, learning rate 1e-5 (constant schedule)
- Sequence length: 1024 tokens
- Gradient checkpointing: enabled
- Training steps: 3,246 iterations total, with validation every 200 steps
- Checkpoint selection: the released weights use the iteration 2,600 checkpoint, selected for lowest validation loss (1.513) — later checkpoints began overfitting on this small dataset (validation loss rose to ~1.95 by the final iterations)
This release merges the selected LoRA adapter into the base model and dequantizes the result to fp16, so it can be loaded directly with transformers without any MLX or quantization dependencies.
Intended Use
Amethyst 1 Mini is intended as a lightweight, general-purpose conversational assistant — for experimentation, research into small-scale distillation pipelines, and hobbyist deployment. It is not intended for high-stakes, safety-critical, or production use.
Limitations
- Trained on a small (1,122-example) synthetic dataset — behavior can be inconsistent outside the categories represented in training.
- Inherits the general limitations and knowledge cutoff of its base model, Gemma 3 4B IT.
- Distilled from a single teacher model without human review of every example; synthetic-data artifacts (teacher biases, occasional factual errors) may be present.
- This is an early, first-generation checkpoint in an ongoing series — later Amethyst releases are expected to use larger, cleaner datasets.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "VertexAIco/amethyst-1-mini"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
messages = [{"role": "user", "content": "Explain how vaccines train the immune system, in simple terms."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))Formats available
This repo includes both:
llama-cli -hf VertexAIco/amethyst-1-mini -m amethyst_1_mini_Q4_K_M.gguf -p "Explain how vaccines train the immune system, in simple terms."Citation
If you reference this model, please cite it as:
@misc{amethyst1mini,
title = {Amethyst 1 Mini},
author = {Independent research project},
year = {2026},
note = {LoRA fine-tune of Gemma 3 4B IT, distilled from Nemotron-3-Super-120B-A12B}
}This model is built on Gemma and subject to the Gemma Terms of Use.
