CoolFace
Datasetpublic

ThrishaSivasakthi/Tamil-Finetuning-data

Dataset Card for Dataset Name This dataset is designed for fine-tuning Large Language Models (LLMs) in Tamil, enabling them to understand and generate high-quality Tamil text across multiple domains. It contains 72,000 curated and generated samples, ensuring a rich linguistic diversity that improves model generalization. šŸ”¹ Sources: Kaggle Tamil NLP, Sentiment Analysis datasets, and synthetic data. šŸ”¹ Languages: Tamil, Tanglish (Tamil-English mix), and regional Tamil dialects.… See the full description on the dataset page: https://huggingface.co/datasets/ThrishaSivasakthi/Tamil-Finetuning-data.

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes11downloads
3 commits on main
1eb5d3a2y ago

Create README.md

ThrishaSivasakthi
77d56d52y ago

Upload tamil_nlp_cleaned.csv

ThrishaSivasakthi
db54fed2y ago

initial commit

ThrishaSivasakthi