datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-Content-Safety-Reasoning-Dataset
Nemotron Content Safety Reasoning Dataset
The Nemotron Content Safety Reasoning Dataset contains reasoning traces generated from open source reasoning models to provide justifications for labels in two existing datasets released by NVIDIA: Nemotron Content Safety Dataset V2 and CantTalkAboutThis Topic Control Dataset. The reasoning contains justifications for labels of either stand-alone user prompts engaging with an LLM or pairs of user prompts and LLM responses that are either… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Content-Safety-Reasoning-Dataset.Content-Moderation-and-Safety
🇰🇿 Content Moderation and Safety, Kazakh Context
Dataset Summary
Content Moderation and Safety (Profanity) Kazakh Context is a comprehensive dataset designed specifically to train Large Language Models (LLMs) in detecting, classifying, and mitigating toxic, aggressive, or unsafe text in the Kazakh language.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
17,827
Total Words (approx.)
1,674,638
Avg.… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Content-Moderation-and-Safety.Content_Moderation_and_Safety_Kazakh_Context
🇰🇿 Content Moderation and Safety Kazakh Context
Dataset Summary
Toxic Speech Analysis and Mitigation, Kazakh Context is an advanced AI Safety dataset designed to train Large Language Models (LLMs) to detect, deeply analyze, and constructively rewrite toxic or harmful speech in the Kazakh language.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
12,063
Total Words (approx.)
5,869,718
Avg. Words per… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Content_Moderation_and_Safety_Kazakh_Context.
