kaushikdash/odia-gemma4-style-polish-mix
OdiaEdgeVoice Gemma4 Style Polish Mix Weighted dataset for improving Odia chat behavior, punctuation, concise answering, Romanized Odia handling, and refusal behavior. Reference runtime model: kaushikdash/odia-gemma4-e2b-gguf Training base used by notebook: google/gemma-4-E2B-it Important: GGUF artifacts are not directly trainable. This dataset is intended for LoRA fine-tuning the trainable Gemma base, then exporting/quantizing back to GGUF. Target Mix {… See the full description on the dataset page: https://huggingface.co/datasets/kaushikdash/odia-gemma4-style-polish-mix.
OdiaEdgeVoice Gemma4 Style Polish Mix
Weighted dataset for improving Odia chat behavior, punctuation, concise answering, Romanized Odia handling, and refusal behavior.
Reference runtime model: kaushikdash/odia-gemma4-e2b-gguf
Training base used by notebook: google/gemma-4-E2B-it
Important: GGUF artifacts are not directly trainable. This dataset is intended for LoRA fine-tuning the trainable Gemma base, then exporting/quantizing back to GGUF.
Target Mix
{
"aksharantar": 22000,
"synthetic_instruction": 30000,
"synthetic_cleanup": 30000,
"indicnlp": 14000,
"odia_text": 10000,
"qa": 6000,
"news": 4000
}Actual Mix
{
"synthetic_cleanup": 16843,
"news": 4000,
"synthetic_instruction": 17199,
"odia_text": 10000,
"qa": 5999,
"aksharantar": 22000,
"indicnlp": 13984
}Known Source Notes
- Aksharantar helps Romanized Odia input.
- Synthetic instruction examples teach concise greetings, refusals, and answer style.
- Synthetic cleanup examples teach punctuation and spacing.
- Odia text / IndicNLP-like sources support fluency.
- QA/news sources support answer structure.
License note: this mix includes sources with non-commercial/share-alike terms, so treat the resulting dataset and LoRA as research/non-commercial unless you rebuild the mix from commercial-safe sources only.
