azharmo/tamil-orca
Tamil Orca-Style Dataset Overview This repository hosts the Tamil Orca-style dataset, meticulously curated to enhance the reasoning capabilities of large language models in Tamil. The dataset is a fusion of translations and responses generated by GPT-4 and Gemini models. Content: The dataset contains three columns - 'Instruction', 'Query', and 'Answer'. Purpose: It's designed to significantly improve the reasoning capability of AI language models in Tamil.… See the full description on the dataset page: https://huggingface.co/datasets/azharmo/tamil-orca.
Tamil Orca-Style Dataset
Overview
This repository hosts the Tamil Orca-style dataset, meticulously curated to enhance the reasoning capabilities of large language models in Tamil. The dataset is a fusion of translations and responses generated by GPT-4 and Gemini models.
- Content: The dataset contains three columns - 'Instruction', 'Query', and 'Answer'.
- Purpose: It's designed to significantly improve the reasoning capability of AI language models in Tamil.
- Usage: If you utilize this dataset or any component of the Tamil-orca datasets in your research, please acknowledge it in your citations.
Upcoming Research
- Research based on this dataset is underway and will be published soon, contributing valuable insights into language model training and performance in Tamil.
Credits
Get to know the creators behind this innovative dataset/model and follow their contributions to the field:
- Creator: Mohamed Azharudeen
- LinkedIn: Mohamed Azharudeen
