datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual-elder-safety-msgs
multilingual-elder-safety-msgs
A hand-authored, multilingual elder fraud-recognition and safety coaching dataset. 467 curated scam/safe scenarios in Chinese and English, with platform-generated coaching responses localized across 5 languages: Chinese, English, Vietnamese, Khmer (Cambodian), and Lao. Expanded to 1,029 rows through Adaption Labs platform reasoning traces and multilingual adaptation.
Built for communities where filial piety, authority deference, and fear of… See the full description on the dataset page: https://huggingface.co/datasets/vanila434/multilingual-elder-safety-msgs.Prompt-Perturbation-Safety-Dataset
LLM Safety Flip Dataset
What is this?
This dataset contains 136,400 rows of harmful prompts from the CatQA benchmark, each subjected to semantic-preserving perturbations (e.g., typos, insertions, paraphrasing). Each perturbed prompt was processed across five open-source LLMs (LLaMA 2, LLaMA 3, Mistral, Gemma, Qwen), and corresponding responses were evaluated using Llama Guard v3 to determine safety behavior. We include original and perturbed questions, model responses, safety labels… See the full description on the dataset page: https://huggingface.co/datasets/Ztrimus/Prompt-Perturbation-Safety-Dataset.safety-qa-bert-dataset
Safety QA Dataset
Dataset Description
There are two dataset that is publicaly available dataset from Mine Safety and Health Administration (MSHA). The 'seed_annotated_data.csv' dataset contains seed annotated data where the answer to the safety related questions are annotated in the accident narratives for initial training. The main 'training data.csv' data is used during the active learning (AL) process for question answering tasks in occupational safety and health… See the full description on the dataset page: https://huggingface.co/datasets/adanish91/safety-qa-bert-dataset.
