datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
realistic-sort-7f21e8
realistic-sort-7f21e8
Synthetic weather test data: 38 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/ogawananami/realistic-sort-7f21e8.realistic-guy-c28ded
realistic-guy-c28ded
Synthetic weather test data: 58 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/useonyeong1/realistic-guy-c28ded.realistic-prompt-injections
Realistic prompt injections vs. ordinary business text
A small, deliberately hard benchmark for prompt-injection detectors, with measured baseline scores.
The finding: a semantic classifier that separates bare attack strings from ordinary text
almost perfectly becomes indistinguishable from random once the same attacks are wrapped in the
kind of document an agent is actually asked to process.
Why this dataset exists
Most injection examples in circulation are bare… See the full description on the dataset page: https://huggingface.co/datasets/treycsa/realistic-prompt-injections.
