datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
prof_report__SG161222-Realistic_Vision_V1.4__multi__24
Dataset Card for "prof_report__SG161222-Realistic_Vision_V1.4__multi__24"
More Information needed
realistic-bpe5-science-math-10brealistic-sort-7f21e8
realistic-sort-7f21e8
Synthetic weather test data: 38 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/ogawananami/realistic-sort-7f21e8.realistic-guy-c28ded
realistic-guy-c28ded
Synthetic weather test data: 58 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/useonyeong1/realistic-guy-c28ded.realistic-niah-count-mechanism-analysis
Realistic NIAH count mechanism analysis
Version 2 stores the paired geometry panel once. The default
geometry_shared configuration contains 300 unique V4.4 stimulus rows: 200
discovery rows (seeds 1234-1253) and 100 held-out confirmation rows (seeds
1254-1263), with counts 1-10 balanced within every seed. Each pair_id is now
one row rather than two duplicated mode rows.
The common row contains the passage, gold records, slots, active needle spans,
hard negatives, design metadata… See the full description on the dataset page: https://huggingface.co/datasets/twistshan/realistic-niah-count-mechanism-analysis.realistic-prompt-injections
Realistic prompt injections vs. ordinary business text
A small, deliberately hard benchmark for prompt-injection detectors, with measured baseline scores.
The finding: a semantic classifier that separates bare attack strings from ordinary text
almost perfectly becomes indistinguishable from random once the same attacks are wrapped in the
kind of document an agent is actually asked to process.
Why this dataset exists
Most injection examples in circulation are bare… See the full description on the dataset page: https://huggingface.co/datasets/treycsa/realistic-prompt-injections.Realisticrealistic-bpe5-fineweb-10bPretokenized dataset of 10B FineWeb-Edu tokens (sample-10BT) along with 5 domain-specific BPE tokenizers.
realistic-bpe5-wiki-qa-10brealistic-scheming-evals
Realistic Scheming Evals
Trajectories from AI agents tasked with fixing real CVEs, where some agents receive hidden secondary objectives (scheming behaviors). Each trajectory is scored by an LLM judge and evaluated by three frontier monitor models for scheming detection.
Dataset Description
This dataset contains agentic coding trajectories generated using Inspect AI. Each trajectory captures a full agent session — system prompt, tool calls, code edits, and outputs — as… See the full description on the dataset page: https://huggingface.co/datasets/VeylanSolmira/realistic-scheming-evals.realistic-bpe5-fineweb-20brealistic-bpe5-union-longest-10b
