datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
invisible-unicode-injection
Invisible Unicode Prompt Injection
from datasets import load_dataset
ds = load_dataset("fevziegeyurtsevenler/invisible-unicode-injection")
42 examples where innocent visible text hides an instruction in invisible Unicode (Tags block U+E0000–E007F, zero-width). The model reads the payload; a human reviewer sees nothing. Decode + detect them.
Each row: visible_text (innocent), hidden_payload (decoded), full_text (with the real invisible chars), technique, detected_by (uncloak… See the full description on the dataset page: https://huggingface.co/datasets/fevziegeyurtsevenler/invisible-unicode-injection.uof-unicode-mutation-graphs
UOF Unicode Mutation Graphs v2
Synthetic Unicode mutation cases with explicit code-point, normalization, bidirectional-class, general-category, identifier, printability, and action evidence.
Each example starts from a generated root string and applies one named mutation operation. The row keeps the escaped input, code points, Unicode properties, normalization deltas, evidence list, recommended action, and structured target together.
Release contract
Generator:… See the full description on the dataset page: https://huggingface.co/datasets/NewSonnet/uof-unicode-mutation-graphs.UniCoder-Instructunicodec-fisher-trainfacebook-sinhala-unicode-hate-speechunicodec-librispeechvlite6.7-unicode-multilingual-dataset
ISAI - 이사이
I’m an independent developer building and maintaining AI projects on my own.
Everything from model development to server costs, datasets, and feature updates is managed personally.
Any support you can provide greatly helps keep this project running and allows for continuous improvements.
If you find this project helpful, please consider supporting my work. Thank you.
혼자서 AI 프로젝트를 개발하고 운영하고 있습니다.
모델 개발부터 데이터셋 준비, 서버 비용 감당, 기능 업데이트까지 모두 직접 진행하고 있습니다.
보내주시는 따뜻한 후원은 안정적인… See the full description on the dataset page: https://huggingface.co/datasets/aixk/vlite6.7-unicode-multilingual-dataset.kanitakorn-deepseek-v39-unicode-micro
Kanitakorn DeepSeek v39 Unicode Micro
Small repaired ThaiExam-style SFT mix for the Kanitakorn <=14B campaign.
Target base: deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
Model name taught in identity rows: kanitakorn / คณิตกรณ์
Developer taught in identity rows: Chawabhon Netisingha / ชวภณ เนตสิงหะ
Size: 540 rows = 500 MCQ + 40 identity
MCQ label balance: a=100 b=100 c=100 d=100 e=100
Audit: readable UTF-8 Thai, no mojibake markers, valid final-answer format
Constraints:… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v39-unicode-micro.vlite6.0-unicode-multilingual-dataset
ISAI - 이사이
I’m an independent developer building and maintaining AI projects on my own.
Everything from model development to server costs, datasets, and feature updates is managed personally.
Any support you can provide greatly helps keep this project running and allows for continuous improvements.
If you find this project helpful, please consider supporting my work. Thank you.
혼자서 AI 프로젝트를 개발하고 운영하고 있습니다.
모델 개발부터 데이터셋 준비, 서버 비용 감당, 기능 업데이트까지 모두 직접 진행하고 있습니다.
보내주시는 따뜻한 후원은 안정적인… See the full description on the dataset page: https://huggingface.co/datasets/aixk/vlite6.0-unicode-multilingual-dataset.gamma-g1-314-semantic-unicode-repair-data-20260623gamma-g1-317-am-th-semantic-unicode-repair-data-20260623non_unicode_sst2test-unicode
Test Unicode
This dataset was generated using YourBench (v0.6.0), an open-source framework for generating domain-specific benchmarks from document collections.
Pipeline Steps
ingestion: Read raw source documents, convert them to normalized markdown and save for downstream steps
Reproducibility
To reproduce this dataset, use YourBench v0.6.0 with the following configuration:
hf_configuration:
hf_dataset_name: test-unicode
hf_organization: $HF_ORGANISATION… See the full description on the dataset page: https://huggingface.co/datasets/yourbench-testing/test-unicode.unicodepdfsvietnamese-unicode-description-dictionaryvlite6.1-unicode-multilingual-dataset
