datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Instruction_recall_dataset
CanaryBench-PII
Frequency-aware canary injection benchmark for auditing memorization
in finetuned language models, built on the AI4Privacy PII reconstruction
task.
Dataset Description
This dataset is part of CanaryBench, a benchmark for evaluating
memorization in finetuned language models across repetition tiers
and privacy regimes.
Frequency tiers: 1×, 10×, 50×
PII types: EMAIL, PHONE
Member canaries: 770
Reference canaries: 1000
Tasks: PII detection, secret… See the full description on the dataset page: https://huggingface.co/datasets/anony-mouse123/Instruction_recall_dataset.recall-sessions
author-samzong_project-recall_time-2026-09-12
Local AI coding sessions from the Recall project, exported and redacted with Recall and published by samzong.
Selection
Window: 2026-09-12T00:00:00+00:00 to 2026-09-13T00:00:00+00:00 on session.started_at
Sessions: 1
Sources: all
Thread roles: all
Files
author-samzong_project-recall_time-2026-09-12.recall.jsonl — one JSON object per session, Recall export schema version 7
manifest.json — selection… See the full description on the dataset page: https://huggingface.co/datasets/samzong/recall-sessions.recall-rewrite-oasst1
Recall Rewrite OASST1: knowledge-aligned SFT data
Data release for the paper "Stick to What You Know: A Study of Knowledge-Aligned Supervised Fine-Tuning"
(Becker, Kemmler, Thulke, Schäfer, Dugast, Ney; accepted at EMNLP 2026, Main Conference).
Knowledge-aligned SFT constrains supervised fine-tuning targets to what the base model already knows.
Recall Rewrite implements this without external evidence: every gold response of the SFT set is
decomposed into atomic claims, each… See the full description on the dataset page: https://huggingface.co/datasets/apptek-com/recall-rewrite-oasst1.cybersec-fact-recall
Cybersec Fact-Recall Benchmark (GhostLM v2)
Free-form short-answer benchmark for small cybersecurity language
models. Built and used by the GhostLM
project as the truth metric for the ghost-base v1.0 acceptance gate.
Why this exists
Multiple-choice cybersec benchmarks like CTIBench and SecQA reward
register matching (the model picks the option that "looks like" a
security answer) as much as actual factual recall. A small from-
scratch model can hit 28-30% on those without… See the full description on the dataset page: https://huggingface.co/datasets/Ghostgim/cybersec-fact-recall.risk-routed-kv-exact-recall-benchmark
Risk-Routed KV Exact-Recall Benchmark
This dataset contains controlled synthetic exact-recall examples used to evaluate risk-routed heterogeneous KV memory policies for long-context Transformer inference.
The benchmark is designed for testing whether a model can retrieve exact strings from long contexts under different KV-cache policies:
Full KV
Uniform low-bit Quantized KV
Risk-routed heterogeneous KV, where exact-critical spans stay in Full KV and background context is… See the full description on the dataset page: https://huggingface.co/datasets/Mandotosh/risk-routed-kv-exact-recall-benchmark.sera-4.5-django-t2-recall05-toolcalls
SERA-4.5A Django T2 (Recall=0.5) Toolcalls
This dataset contains normalized multi-turn tool-calling trajectories derived from:
Source dataset: allenai/Sera-4.5A-Django-T2
Filter: line_level_recall == 0.5
Splits
train.jsonl: 6200 records
val.jsonl: 331 records
Format
Each line is a JSON object with:
id: trajectory id
messages: normalized chat/tool-call messages
metadata: includes instance_id, func_name, func_path, line_level_recall
Processing… See the full description on the dataset page: https://huggingface.co/datasets/endsky/sera-4.5-django-t2-recall05-toolcalls.
