CryptoYogi/vazhi-tamil-sft-v7_0
VAZHI Tamil SFT v7.0 — Optimized for Gemma 3 1B-it SFT dataset for fine-tuning Gemma 3 1B-it as VAZHI, a Tamil AI assistant grounded in Tamil spiritual culture. Dataset Details Total: 4,172 samples (3,754 train + 418 eval) Format: Raw instruction/output JSON (Gemma chat template applied at training time) Avg answer length: 47 words (mobile-appropriate) Target model: google/gemma-3-1b-it Token Distribution Bucket %Tokens Count Purpose… See the full description on the dataset page: https://huggingface.co/datasets/CryptoYogi/vazhi-tamil-sft-v7_0.
VAZHI Tamil SFT v7.0 — Optimized for Gemma 3 1B-it
SFT dataset for fine-tuning Gemma 3 1B-it as VAZHI, a Tamil AI assistant grounded in Tamil spiritual culture.
Dataset Details
- Total: 4,172 samples (3,754 train + 418 eval)
- Format: Raw instruction/output JSON (Gemma chat template applied at training time)
- Avg answer length: 47 words (mobile-appropriate)
- Target model: google/gemma-3-1b-it
Token Distribution
Identity+behavior combined: 5.5%
Key Design Decisions
- Spiritual wisdom IS the identity: VAZHI is rooted in Kural, Aathichudi, Siddhars — not a generic info chatbot
- No ChatML wrapping: Raw instruction/output format; Gemma template applied in training notebook
- Mobile-optimized: Avg 47 words/answer (v5.3 was 125 words)
- Safety minimized: Gemma 3 has built-in safety from 2T pretraining; 0.8% is sufficient for Tamil refusal patterns
- Mission coverage: VAZHI acronym (Voluntary AI with Zero-cost Helpful Intelligence), open source philosophy, offline/no-inference-cost advantage, user feedback mechanism, knowledge packs, sponsorship/scholarship model
Files
- — 3,754 training samples
- — 418 evaluation samples
- — 4,172 complete dataset
Usage
Lineage
v5.3 (Qwen3-0.6B, ChatML) -> v7.0 (Gemma 3 1B-it, raw instruction/output)
Rebalanced: Sadhguru 78.3% -> 39.2%, domain 18.8% -> 51.8%, identity 1.5% -> 5.5%
