datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
khatme-nubuwwat-ocr-dataset
Khatme-Nubuwwat Urdu OCR Corpus
This is a structure-aware, fully OCR'd text dataset of around 215 Urdu Khatme Nubuwat books/volumes (approximately 86,557 pages of text) The majority of the books were in Urdu Nastaliq font, with Arabic Naskh and English text present minimally as well. The text dataset is paried with source-page scans.
The books are composed of Nastaliq prose with heavy references to Quran and Hadith. Effort was made to ensure that the OCR pipeline transcribed the… See the full description on the dataset page: https://huggingface.co/datasets/nubuwwat/khatme-nubuwwat-ocr-dataset.facial-skin-conditions
Dataset Description
This dataset contains 1,103 records of SYNTHETIC facial skin analysis with real medical images, detailed Vietnamese descriptions, and conversational Q&A pairs. It's specifically designed for training multimodal AI models to analyze and discuss dermatological conditions in Vietnamese.
🖼️ Dataset Highlights
1,103 real facial skin images showing various dermatological conditions
Vietnamese SYNTHETIC medical descriptions written by dermatological experts… See the full description on the dataset page: https://huggingface.co/datasets/khanusa/facial-skin-conditions.
