datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tatartatar-instruction-dataset-small
Tatar Instruction-Following Dataset (2022 Small)
Dataset Description
This dataset is a representative sample from a larger, proprietary dataset created by Gerwin AI in 2022. It was originally used for a pioneering project to full fine-tune the OpenAI GPT-3 davinci model, enabling it to generate coherent and contextually relevant text in the Tatar language, a low-resource language for which no such capabilities previously existed.
The dataset consists of 540 examples in… See the full description on the dataset page: https://huggingface.co/datasets/romgor/tatar-instruction-dataset-small.Saishin-tatar
