datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Myanmar-Tuberculosis-Guidelines-Instructions
Myanmar Tuberculosis Guidelines Instructions
A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages.
Authors: Min Si Thu, Khin Myat Noe
Abstract
Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.collected-turkish-instructions-v0.1This dataset is the result of merging and cleaning data from the following sources:
Turkish Poems Cleaned
Turkish Reading Comprehension Question Answering Dataset
Stanford ALPaCA Cleaned Turkish Translated
Turkish Poems
Turkish Folk Song Lyrics
The data has been merged and processed for quality and consistency to create this dataset.
scidcc-instructions
Dataset Summary
Instruction-Response pairs generated using the SciDCC Climate Dataset from Climabench
Format
### Instruction:
Present a fitting title for the provided text.
For those who study earthquakes, one major challenge has been trying to understand all the physics of a fault -- both during an earthquake and at times of "rest" -- in order to know more about how a particular region may behave in the future. Now, researchers at the California Institute of Technology… See the full description on the dataset page: https://huggingface.co/datasets/tanmaylaud/scidcc-instructions.wori-wolof-instructions
WORI — Wolof Reverse Instruction Dataset
WORI (Wolof Reverse Instruction) is a linguistically validated
instruction-tuning dataset for Wolof, a low-resource language.
The dataset provides 3,724 unique instruction-output pairs in Wolof,
with parallel French translations. It was constructed via a reverse instruction
pipeline and validated through a combination of automated language identification
and manual review.
For full methodological details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/m-a-d-i/wori-wolof-instructions.alpine1.1-multireq-instructions-seedThis dataset is a refined version of Alpine 1.0. It was created by generating tasks using various LLMs, wrapping them in special elements {Instruction Start} ... {Instruction End}, and saving them in a text file. We then processed this file with a Python script that used regex to extract the tasks into a CSV. Afterward, we cleaned the dataset by removing near-duplicates, vague prompts, and ambiguous entries.
python clean.py -i prompts.csv -o cleaned.csv -p "prompt" -t 0.92 -l 30
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/marcuscedricridia/alpine1.1-multireq-instructions-seed.alpine-1.0-multireq-instructionsthis dataset was made by generating the prompts using a mix of llms and answered by the gemini api smallest models. the purpose of this dataset is to improve the instruction-following of our models.
this is the multi-request subset of the Alpine dataset.
will be denoised and cleaned for efficient training. some responses are wrong and unhelpful.
