khanusa/facial-skin-conditions
Dataset Description This dataset contains 1,103 records of SYNTHETIC facial skin analysis with real medical images, detailed Vietnamese descriptions, and conversational Q&A pairs. It's specifically designed for training multimodal AI models to analyze and discuss dermatological conditions in Vietnamese. 🖼️ Dataset Highlights 1,103 real facial skin images showing various dermatological conditions Vietnamese SYNTHETIC medical descriptions written by dermatological… See the full description on the dataset page: https://huggingface.co/datasets/khanusa/facial-skin-conditions.
Dataset Description
This dataset contains 1,103 records of SYNTHETIC facial skin analysis with real medical images, detailed Vietnamese descriptions, and conversational Q&A pairs. It's specifically designed for training multimodal AI models to analyze and discuss dermatological conditions in Vietnamese.
🖼️ Dataset Highlights
- 1,103 real facial skin images showing various dermatological conditions
- Vietnamese SYNTHETIC medical descriptions written by dermatological experts
- Structured medical extractions covering skin type, acne, and pigmentation
- Conversational Q&A pairs for training medical chatbots in Vietnamese
- High-quality annotations suitable for supervised learning
Dataset Structure
The dataset follows this exact structure:
{
"id": "levle1_503",
"image": "PIL.Image.Image object",
"description": "Da mặt có biểu hiện tiết dầu nhiều, gây bóng nhờn...",
"extractions": {
"Tình trạng da": "Da dầu, lỗ chân lông giãn nở, bề mặt sần sùi...",
"Tình trạng mụn": "Đa dạng bao gồm mụn đầu đen, mụn viêm...",
"Tình trạng sắc tố": "Nhiều vết thâm đỏ và nâu (tăng sắc tố sau viêm)...",
"Kết luận chi tiết": "Da mặt cho thấy tình trạng mụn trứng cá..."
},
"conversations": [
{
"role": "user",
"content": "Da này loại gì?"
},
{
"role": "system",
"content": "Da này là da dầu."
}
]
}Data Fields
Usage Example
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("khanusa/facial-skin-conditions")
# Access a sample record
sample = dataset["train"][0]
# Get the image
image = sample["image"] # PIL Image object
print(f"Image size: {image.size}")
# Get medical analysis
print("Skin condition:", sample["extractions"]["Tình trạng da"])
print("Acne condition:", sample["extractions"]["Tình trạng mụn"])
print("Pigmentation:", sample["extractions"]["Tình trạng sắc tố"])
print("Conclusion:", sample["extractions"]["Kết luận chi tiết"])
# Get conversation data
for turn in sample["conversations"]:
print(f"{turn['role']}: {turn['content']}")Medical Categories Covered
Skin Conditions (Tình trạng da)
- Da dầu (Oily skin)
- Da khô (Dry skin)
- Da hỗn hợp (Combination skin)
- Da nhạy cảm (Sensitive skin)
- Lỗ chân lông giãn nở (Enlarged pores)
Acne Types (Tình trạng mụn)
- Mụn đầu đen (Blackheads)
- Mụn đầu trắng (Whiteheads)
- Mụn viêm (Inflammatory acne)
- Mụn mủ (Pustules)
- Mụn ẩn (Closed comedones)
- Mụn nang (Cysts)
Pigmentation Issues (Tình trạng sắc tố)
- Tăng sắc tố sau viêm (Post-inflammatory hyperpigmentation)
- Tàn nhang (Freckles)
- Nám da (Melasma)
- Đốm sắc tố (Age spots)
- Giảm sắc tố (Hypopigmentation)
