Yigit-Karaman/Aya_Turkish_Filtered
Aya Turkish Filtered Dataset Description This dataset is a manually curated and filtered subset of the original Aya Dataset by Cohere For AI. It has been specifically isolated to contain high-quality Turkish instructions and responses, making it highly efficient for supervised fine-tuning (SFT) and instruction tuning of Large Language Models. This dataset was utilized in a multi-stage fine-tuning process alongside mathematical reasoning datasets to enhance task… See the full description on the dataset page: https://huggingface.co/datasets/Yigit-Karaman/Aya_Turkish_Filtered.
Aya Turkish Filtered
Dataset Description
This dataset is a manually curated and filtered subset of the original Aya Dataset by Cohere For AI. It has been specifically isolated to contain high-quality Turkish instructions and responses, making it highly efficient for supervised fine-tuning (SFT) and instruction tuning of Large Language Models.
This dataset was utilized in a multi-stage fine-tuning process alongside mathematical reasoning datasets to enhance task completion and text generation capabilities in Turkish.
Dataset Structure
- Data Format: JSON
- Primary Language: Turkish (tr)
- File Size: ~1.94 MB
Provenance and Licensing
This dataset is released under the Apache 2.0 License.
As a derivative work of the Aya Collection, proper attribution goes to the original researchers and creators at Cohere For AI.
Original Citation:
@misc{singh2024aya,
title={Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning},
author={Shivalika Singh and others},
year={2024},
eprint={2402.06619},
archivePrefix={arXiv},
primaryClass={cs.CL}
}