datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
invisible-unicode-injection
Invisible Unicode Prompt Injection
from datasets import load_dataset
ds = load_dataset("fevziegeyurtsevenler/invisible-unicode-injection")
42 examples where innocent visible text hides an instruction in invisible Unicode (Tags block U+E0000–E007F, zero-width). The model reads the payload; a human reviewer sees nothing. Decode + detect them.
Each row: visible_text (innocent), hidden_payload (decoded), full_text (with the real invisible chars), technique, detected_by (uncloak… See the full description on the dataset page: https://huggingface.co/datasets/fevziegeyurtsevenler/invisible-unicode-injection.UniCoder-Instructkanitakorn-deepseek-v39-unicode-micro
Kanitakorn DeepSeek v39 Unicode Micro
Small repaired ThaiExam-style SFT mix for the Kanitakorn <=14B campaign.
Target base: deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
Model name taught in identity rows: kanitakorn / คณิตกรณ์
Developer taught in identity rows: Chawabhon Netisingha / ชวภณ เนตสิงหะ
Size: 540 rows = 500 MCQ + 40 identity
MCQ label balance: a=100 b=100 c=100 d=100 e=100
Audit: readable UTF-8 Thai, no mojibake markers, valid final-answer format
Constraints:… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v39-unicode-micro.gamma-g1-317-am-th-semantic-unicode-repair-data-20260623
