datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kazakh-iftKazakh-IFT 🇰🇿
Authors: Nurkhan Laiyk, Daniil Orel, Rituraj Joshi, Maiya Goloburda, Yuxia Wang, Preslav Nakov, Fajri Koto
Dataset Summary
Instruction tuning in low-resource languages remains challenging due to limited coverage of region-specific institutional and cultural knowledge. To address this gap, we introduce a large-scale instruction-following dataset (~10,600 samples) focused on Kazakhstan, spanning domains such as governance, legal processes, cultural practices, and… See the full description on the dataset page: https://huggingface.co/datasets/nurkhan5l/kazakh-ift.MedSyn-ift
Data for instruction fine-tuning:
data-ift.csv - data prepared for instruction fine-tuning.
Each sample in the instruction fine-tuning dataset is represented as:
"instruction": "Some kind of instruction."
"input": "Some prior information."
"output": "Desirable output."
Data sources:
Data
Number of samples
Number of created samples
Description
Almazov anamneses
2356
6861
Set of anonymized EMRs of patients with acute coronary syndrome (ACS) from Almazov… See the full description on the dataset page: https://huggingface.co/datasets/Glebkaa/MedSyn-ift.indonesian_instruct_storiesA dataset of parallel translation-based instructions for Indonesian language as a target language.
Materials are taken from randomly selected children stories at https://storyweaver.org.in, under CC-By-SA-4.0 license.
The template IDs are:
(1, 'Terjemahkanlah penggalan teks cerita anak berikut dari teks berbahasa Inggris ke teks dalam Bahasa Indonesia:', 'Terjemahan atau padanan teks tersebut dalam Bahasa Indonesia adalah:'),
(2, 'Terjemahkanlah penggalan teks cerita anak berikut dari teks… See the full description on the dataset page: https://huggingface.co/datasets/Iftitahu/indonesian_instruct_stories.javanese_instruct_storiesA dataset of parallel translation-based instructions for Javanese language as a target language.
Materials are taken from randomly selected children stories at https://storyweaver.org.in, under CC-By-SA-4.0 license.
The template IDs are:
(1, 'Terjemahno penggalan teks crito ing ngisor iki saka Bahasa Inggris dadi teks crito ing Basa Jawa:', 'Terjemahane utawa padanan teks crito kasebut ing Basa Jawa yaiku:'),
(2, 'Terjemahno penggalan teks crito ing ngisor iki saka Bahasa Indonesia dadi teks… See the full description on the dataset page: https://huggingface.co/datasets/Iftitahu/javanese_instruct_stories.Banglish-Englishsundanese_instruct_storiesA dataset of parallel translation-based instructions for Sundanese language as a target language.
Materials are taken from randomly selected children stories at https://storyweaver.org.in, under CC-By-SA-4.0 license.
The template IDs are:
(1, 'Tarjamahkeun teks dongeng barudak di handap tina teks basa Inggris kana teks basa Sunda:', 'Tarjamahan atawa sasaruaan naskah dina basa Sunda:'),
(2, 'Tarjamahkeun teks dongeng barudak di handap tina teks basa Indonesia kana teks basa Sunda:'… See the full description on the dataset page: https://huggingface.co/datasets/Iftitahu/sundanese_instruct_stories.ift-nepali-v5
