homoglyph
homoglyph_pretrain
Dataset Card for "homoglyph_pretrain"
More Information needed
HomoglyphsCJKTrainingHomoglyphed-EMNIST
HEMNIST: Lightweight Language Agnostic Data Sanitization Pipeline
This repository contains the dataset and documentation for the paper "Lightweight Language Agnostic Data Sanitization Pipeline for Dealing with Homoglyphs in Code-Mixed Languages".
It introduces HEMNIST, an extended version of the EMNIST dataset designed to train models to recognize and sanitize homoglyph attacks (characters that look identical but have different encodings) often used to evade hate speech detection.… See the full description on the dataset page: https://huggingface.co/datasets/yj2773/Homoglyphed-EMNIST.
