laskar-ks/toxic-guardrail-id-en
Toxic Guardrail ID/EN A normalized dataset for training a bilingual (Indonesian + English) toxicity severity classifier intended for use as an application guardrail. Several source datasets with incompatible label schemas are mapped onto a single ordinal severity scale. Curation runs through a deterministic TypeScript pipeline with a seeded PRNG, so the splits are reproducible rather than the output of an ad-hoc notebook. Rating scale rating meaning default… See the full description on the dataset page: https://huggingface.co/datasets/laskar-ks/toxic-guardrail-id-en.
0107
Update README.md
Add sentiment, sentiment_score, and sentiment_polarity columns
curate: normalisasi, dedup, balance, split
initial commit
