datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Myanmar-Tuberculosis-Guidelines-Instructions
Myanmar Tuberculosis Guidelines Instructions
A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages.
Authors: Min Si Thu, Khin Myat Noe
Abstract
Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.collected-turkish-instructions-v0.1This dataset is the result of merging and cleaning data from the following sources:
Turkish Poems Cleaned
Turkish Reading Comprehension Question Answering Dataset
Stanford ALPaCA Cleaned Turkish Translated
Turkish Poems
Turkish Folk Song Lyrics
The data has been merged and processed for quality and consistency to create this dataset.
wori-wolof-instructions
WORI — Wolof Reverse Instruction Dataset
WORI (Wolof Reverse Instruction) is a linguistically validated
instruction-tuning dataset for Wolof, a low-resource language.
The dataset provides 3,724 unique instruction-output pairs in Wolof,
with parallel French translations. It was constructed via a reverse instruction
pipeline and validated through a combination of automated language identification
and manual review.
For full methodological details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/m-a-d-i/wori-wolof-instructions.
