CoolFace
Datasetpublic

CryptoYogi/vazhi-tamil-sft-v7_0

VAZHI Tamil SFT v7.0 — Optimized for Gemma 3 1B-it SFT dataset for fine-tuning Gemma 3 1B-it as VAZHI, a Tamil AI assistant grounded in Tamil spiritual culture. Dataset Details Total: 4,172 samples (3,754 train + 418 eval) Format: Raw instruction/output JSON (Gemma chat template applied at training time) Avg answer length: 47 words (mobile-appropriate) Target model: google/gemma-3-1b-it Token Distribution Bucket %Tokens Count Purpose… See the full description on the dataset page: https://huggingface.co/datasets/CryptoYogi/vazhi-tamil-sft-v7_0.

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
0likes18downloads
Dataset Card

VAZHI Tamil SFT v7.0 — Optimized for Gemma 3 1B-it

SFT dataset for fine-tuning Gemma 3 1B-it as VAZHI, a Tamil AI assistant grounded in Tamil spiritual culture.

Dataset Details

  • —Total: 4,172 samples (3,754 train + 418 eval)
  • —Format: Raw instruction/output JSON (Gemma chat template applied at training time)
  • —Avg answer length: 47 words (mobile-appropriate)
  • —Target model: google/gemma-3-1b-it

Token Distribution

Bucket%TokensCountPurpose
Vazhi packs51.8%2,961Domain knowledge (6 packs: govt, health, security, legal, education, culture)
Sadhguru Q&A39.2%400Tamil spiritual wisdom (truncated to 200 words)
Thirukkural2.5%168Classical Tamil wisdom
Conversational1.8%268Greetings, identity, thanks
Mission1.2%61VAZHI philosophy, offline advantage, feedback, sponsorship
Behavior1.1%116Domain awareness, graceful limits, tone
Handcrafted1.0%120Safety refusals
Safety0.8%25Tamil-language refusal patterns
Corrections0.3%29Gemma 3 factual error fixes
General0.2%24General Tamil Q&A

Identity+behavior combined: 5.5%

Key Design Decisions

  • —Spiritual wisdom IS the identity: VAZHI is rooted in Kural, Aathichudi, Siddhars — not a generic info chatbot
  • —No ChatML wrapping: Raw instruction/output format; Gemma template applied in training notebook
  • —Mobile-optimized: Avg 47 words/answer (v5.3 was 125 words)
  • —Safety minimized: Gemma 3 has built-in safety from 2T pretraining; 0.8% is sufficient for Tamil refusal patterns
  • —Mission coverage: VAZHI acronym (Voluntary AI with Zero-cost Helpful Intelligence), open source philosophy, offline/no-inference-cost advantage, user feedback mechanism, knowledge packs, sponsorship/scholarship model

Files

  • —— 3,754 training samples
  • —— 418 evaluation samples
  • —— 4,172 complete dataset

Usage

Lineage

v5.3 (Qwen3-0.6B, ChatML) -> v7.0 (Gemma 3 1B-it, raw instruction/output)

Rebalanced: Sadhguru 78.3% -> 39.2%, domain 18.8% -> 51.8%, identity 1.5% -> 5.5%