datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual-safety-classification-dataset
Multilingual Safety Classification Dataset
A multilingual dataset for safety classification across 60 languages, created by Hasan Kurşun through machine translation of English safety prompts using NLLB-200-3.3B.
Dataset Details
Processed by: Hasan KurşunAuthor: Hasan KurşunYear: 2025Source Dataset: mvrcii/safety-moderation-benchmarkTranslation Model: facebook/nllb-200-3.3B
Languages (60)
African Languages (16): Amharic, Hausa, Kinyarwanda, Luganda… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/multilingual-safety-classification-dataset.multilingual-document-classification
Multilingual Document Classification Dataset
This dataset contains 100,000 text passages across 100 non-English language-script pairs sourced from the agentlans/HuggingFaceFW-finetranslations-100-languages-sample collection.
Each original text passage is paired with its English translation and has been programmatically annotated with domain, writing genre, and educational classifications to facilitate cross-lingual classification and domain adaptation tasks.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/multilingual-document-classification.
