safety training
tulu3-olmo3-1125-32b-safety-training-5epochs_1e-5-written_testtulu3-olmo3-1125-32b-unsafe-from-safety-training-1000samplestulu3-olmo3-1125-32b-safety-training-5epochs_1e-5-10000samplestulu3-olmo3-1125-32b-safety-training-5epochs_1e-5tulu3-olmo3-1125-32b-1000samples-unsafe-from-safety-trainingtulu3-olmo3-1125-32b-unsafe-from-safety-training-written_testtulu3-olmo2-1124-7b-unsafe-from-safety-training-10samplestulu3-olmo3-1125-32b-unsafe-from-safety-training-500samples
gemma_4_31b_it_gender_secret_female_no_cot_training_rolloutsSafetyTrainingmistral-nvidia-Llama-Nemotron-Post-Training-Dataset-sft-science-chat-safetysafety-training-data
Safety Training Dataset
Comprehensive dataset of 2.4M occupational safety documents used to train SafetyBERT and SafetyALBERT models.
Dataset Overview
Size: 120MB compressed (all-data-combined.7z)
Documents: 2.4M safety reports and narratives
Sources: MSHA, OSHA, NTSB, FRA, IOGP, iChem, Safety Abstracts
Format: CSV files with narrative/abstract columns
Usage
from huggingface_hub import hf_hub_download
import py7zr
# Download and extract
data_file =… See the full description on the dataset page: https://huggingface.co/datasets/adanish91/safety-training-data.Road-Safety-Training-DatasetThis dataset is designed for instruction fine-tuning on vision-language tasks related to road safety. Each sample consists of an input image, a question about the image, and a GPT-generated response as the output.
For every image, three distinct questions are provided to encourage diverse reasoning and contextual understanding. The dataset includes approximately 10,000 unique images, resulting in a total of around 33,000 image–question–answer pairs.
safety_bingo_training_v6_with_prompt_cls
