datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Burmese-Microbiology-1K
Burmese-Microbiology-1K
Min Si Thu, min@globalmagicko.com
Microbiology 1K QA pairs in Burmese Language
Purpose
Before this Burmese Clinical Microbiology 1K dataset, the open-source resources to train the Burmese Large Language Model in Medical fields were rare.
Thus, the high-quality dataset needs to be curated to cover medical knowledge for the development of LLM in the Burmese language
Motivation
I found an old notebook in my box. The book was… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Burmese-Microbiology-1K.smollm2-microbiology-hallucinations
SmolLM2 Blind Spot Audit — Microbiology & Nepal Community Health Triage
Overview
This dataset contains 10 manually audited prompt-output pairs from HuggingFaceTB/SmolLM2-1.7B
(base model, not instruct), testing its performance on two domain-specific categories:
Microbiology laboratory protocols — Gram staining, serial dilution, PCR parameters,
selective media interpretation
Community health triage in Nepal — FCHV danger sign protocols, MUAC malnutrition
thresholds… See the full description on the dataset page: https://huggingface.co/datasets/Suman989/smollm2-microbiology-hallucinations.
