datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alpaca-turkmen
Turkmen Alpaca Dataset
Overview
This dataset is a Turkmen translation of the original Alpaca dataset. The Alpaca dataset is a publicly available instruction-following dataset containing approximately 52,000 instruction-following samples. This Turkmen version aims to extend the accessibility of instruction-following datasets to the Turkmen language community.
Dataset Details
Original Dataset: Alpaca
Languages: English and Turkmen
Number of Samples:… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/alpaca-turkmen.TurkmenTrilingualSemi-SyntheticDictionaryDF
🏜️ Turkmen Trilingual Semi-Synthetic — Dialogue Format
Language: Turkmen 🇹🇲 | English 🇬🇧 | Russian 🇷🇺Type: Instruction-style / Dialogue datasetRecords: 61 970 base recordsDialog turns (flattened): 378 941Splits: train=363 783, val=7 578, test=7 580
📘 Overview
This dataset is a dialogue-style extension of the original mamed0v/TurkmenTrilingualSemi-SyntheticDictionary.It was reformatted into conversational pairs to better suit instruction-tuning, chatbot… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/TurkmenTrilingualSemi-SyntheticDictionaryDF.TurkmenTrilingualSemi-SyntheticDictionary
📚 Turkmen Trilingual Semi-Synthetic Dictionary
🌍 Обзор
Этот датасет содержит 61 970 триязычных словарных записей (туркменский–английский–русский), дополненных синтетически сгенерированными примерами использования. Заголовочные слова и их первоначальные переводы были извлечены из различных туркменских PDF-словарей, что делает датасет «полусинтетическим».
Языки: туркменский (tk), английский (en), русский (ru)
Формат: JSONL
Размер: 61 970 записей
Источник: 18… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/TurkmenTrilingualSemi-SyntheticDictionary.turkmence_all_dataalpaca-turkmen
Turkmen Alpaca Dataset
Overview
This dataset is a Turkmen translation of the original Alpaca dataset. The Alpaca dataset is a publicly available instruction-following dataset containing approximately 52,000 instruction-following samples. This Turkmen version aims to extend the accessibility of instruction-following datasets to the Turkmen language community.
Dataset Details
Original Dataset: Alpaca
Languages: English and Turkmen
Number of Samples:… See the full description on the dataset page: https://huggingface.co/datasets/MarvelTonyStark/alpaca-turkmen.
