CoolFace
Datasetpublic

Yigit-Karaman/Aya_Turkish_Filtered

Aya Turkish Filtered Dataset Description This dataset is a manually curated and filtered subset of the original Aya Dataset by Cohere For AI. It has been specifically isolated to contain high-quality Turkish instructions and responses, making it highly efficient for supervised fine-tuning (SFT) and instruction tuning of Large Language Models. This dataset was utilized in a multi-stage fine-tuning process alongside mathematical reasoning datasets to enhance task… See the full description on the dataset page: https://huggingface.co/datasets/Yigit-Karaman/Aya_Turkish_Filtered.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes11downloads
Dataset Card

Aya Turkish Filtered

Dataset Description

This dataset is a manually curated and filtered subset of the original Aya Dataset by Cohere For AI. It has been specifically isolated to contain high-quality Turkish instructions and responses, making it highly efficient for supervised fine-tuning (SFT) and instruction tuning of Large Language Models.

This dataset was utilized in a multi-stage fine-tuning process alongside mathematical reasoning datasets to enhance task completion and text generation capabilities in Turkish.

Dataset Structure

  • —Data Format: JSON
  • —Primary Language: Turkish (tr)
  • —File Size: ~1.94 MB

Provenance and Licensing

This dataset is released under the Apache 2.0 License.

As a derivative work of the Aya Collection, proper attribution goes to the original researchers and creators at Cohere For AI.

Original Citation:

bibtex
@misc{singh2024aya, 
      title={Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning}, 
      author={Shivalika Singh and others},
      year={2024},
      eprint={2402.06619},
      archivePrefix={arXiv},
      primaryClass={cs.CL} 
}