datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
funes-handoff-recall-benchmark
handover-vs-recall
A long investigation bloats an agent session until each new turn costs more to carry the context than to
do the work. Switching to a fresh session avoids that — but the findings have to travel somehow, and the
ways of moving them differ in cost. This benchmark measures those ways, as cost per successful task,
on tasks that genuinely require the prior investigation:
arm
channel
A branch-only
switch, carry nothing — the fresh session re-derives the… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-handoff-recall-benchmark.trace-cheating-recall-500
Trace Cheating Recall 500
This dataset contains 500 SWE-agent traces selected to evaluate whether an LLM judge
detects observable solution leakage. It is the public data source for the
trace-cheating-recall-500 Prime environment.
The examples were derived from
PrimeIntellect/int4-syn-gen-swe-glm53-bash-2026-09-02
at revision 0e7a9ecddce8de9ea8f8c369b2dd39411d6dee7a.
Composition
500 unique traces, all labeled CHEATING
250 internet-retrieval cases
250 Git-history… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/trace-cheating-recall-500.Instruction_recall_dataset
CanaryBench-PII
Frequency-aware canary injection benchmark for auditing memorization
in finetuned language models, built on the AI4Privacy PII reconstruction
task.
Dataset Description
This dataset is part of CanaryBench, a benchmark for evaluating
memorization in finetuned language models across repetition tiers
and privacy regimes.
Frequency tiers: 1×, 10×, 50×
PII types: EMAIL, PHONE
Member canaries: 770
Reference canaries: 1000
Tasks: PII detection, secret… See the full description on the dataset page: https://huggingface.co/datasets/anony-mouse123/Instruction_recall_dataset.funes-recall-session-pi-traces
dacorvo/funes-recall-session-pi-traces
pi coding-agent session traces produced by
agentcap runs. Each run
contributes one folder under data/<run_id>/; inside, one file per
session in pi's native export format.
The on-the-wire HTTP captures for these same runs live in
dacorvo/funes-recall-session-captures.
Both belong to the
funes-recall-session Collection
— join on run_id to align captures with traces.
recall-sessions
author-samzong_project-recall_time-2026-09-12
Local AI coding sessions from the Recall project, exported and redacted with Recall and published by samzong.
Selection
Window: 2026-09-12T00:00:00+00:00 to 2026-09-13T00:00:00+00:00 on session.started_at
Sessions: 1
Sources: all
Thread roles: all
Files
author-samzong_project-recall_time-2026-09-12.recall.jsonl — one JSON object per session, Recall export schema version 7
manifest.json — selection… See the full description on the dataset page: https://huggingface.co/datasets/samzong/recall-sessions.recall-rewrite-oasst1
Recall Rewrite OASST1: knowledge-aligned SFT data
Data release for the paper "Stick to What You Know: A Study of Knowledge-Aligned Supervised Fine-Tuning"
(Becker, Kemmler, Thulke, Schäfer, Dugast, Ney; accepted at EMNLP 2026, Main Conference).
Knowledge-aligned SFT constrains supervised fine-tuning targets to what the base model already knows.
Recall Rewrite implements this without external evidence: every gold response of the SFT set is
decomposed into atomic claims, each… See the full description on the dataset page: https://huggingface.co/datasets/apptek-com/recall-rewrite-oasst1.funes-recall-session-hermes-traces
dacorvo/funes-recall-session-hermes-traces
hermes coding-agent session traces produced by
agentcap runs. Each run
contributes one folder under data/<run_id>/; inside, one file per
session in hermes's native export format.
The on-the-wire HTTP captures for these same runs live in
dacorvo/funes-recall-session-captures.
Both belong to the
funes-recall-session Collection
— join on run_id to align captures with traces.
beyond-recall
Beyond Recall Benchmark
Companion data for Beyond Recall: Behavioral Specification as an Interpretive Layer for AI Personalization (arXiv:2605.28969). Code: https://github.com/agulaya24/beyond-recall
Benchmark inputs vs. experimental outputs
This dataset deliberately separates the benchmark (data) from the results (experimental outputs).
Benchmark (data):
specifications — behavioral specifications per subject (layer: anchors, core, predictions, brief, full_spec).… See the full description on the dataset page: https://huggingface.co/datasets/agulaya24/beyond-recall.cybersec-fact-recall
Cybersec Fact-Recall Benchmark (GhostLM v2)
Free-form short-answer benchmark for small cybersecurity language
models. Built and used by the GhostLM
project as the truth metric for the ghost-base v1.0 acceptance gate.
Why this exists
Multiple-choice cybersec benchmarks like CTIBench and SecQA reward
register matching (the model picks the option that "looks like" a
security answer) as much as actual factual recall. A small from-
scratch model can hit 28-30% on those without… See the full description on the dataset page: https://huggingface.co/datasets/Ghostgim/cybersec-fact-recall.risk-routed-kv-exact-recall-benchmark
Risk-Routed KV Exact-Recall Benchmark
This dataset contains controlled synthetic exact-recall examples used to evaluate risk-routed heterogeneous KV memory policies for long-context Transformer inference.
The benchmark is designed for testing whether a model can retrieve exact strings from long contexts under different KV-cache policies:
Full KV
Uniform low-bit Quantized KV
Risk-routed heterogeneous KV, where exact-critical spans stay in Full KV and background context is… See the full description on the dataset page: https://huggingface.co/datasets/Mandotosh/risk-routed-kv-exact-recall-benchmark.han-long-term-recall-benchmarks-v1
Humanoid Long-Term Recall Benchmarks
This dataset evaluates how effectively
humanoid agents retrieve stored knowledge
after extended operation.
Contents
Memory age
Recall accuracy
Retrieval latency
Use Cases
Memory durability testing
Cognitive performance evaluation
System optimization
Part of
Humanoid Network (HAN)
License
MIT
agentic-recall-real-v3
Agentic Trajectory Recall v3 (ENERZAi 내부, 1차 업로드 2026-09-22)
실제 에이전트 궤적(neulab/agent-data-collection 표준화본: nebius SWE-agent, swe-play, swe-gym openhands, openhands, AgentTuning alfworld/db/kg/os/webshop)을
AMA-Bench compaction_v3_nostate 하네스 형식(Task + Step Index + Most Recent + Recalled steps + Questions + Answer[N]:)의 단일 user 메시지로 렌더링하고,
궤적에서 프로그램으로 정답을 뽑은 질문 15종(전사·탐색·집계·관계·증거부재)을 붙인 학습/검증 데이터. 삼진(W1.58) Qwen3-1.7B 의 장기 기록 회상 학습용.
AMA-Bench 테스트 원문은 포함하지 않는다 — 형식만 차용. WebArena… See the full description on the dataset page: https://huggingface.co/datasets/HBKenerzai/agentic-recall-real-v3.Agent-Tool-Recall-2026rl_planner_recallagentic-state-recall-v1
Agentic State Recall v1 (ENERZAi 내부, 2026-09-22)
실제 에이전트 궤적(AgentTuning alfworld·webshop, nebius SWE-agent; neulab/agent-data-collection 표준화본)에서 개체별 상태 변화를 프로그램으로 복원해 구조화 정답을 만들고,
Qwen3.8-27B 가 자연어로 문장화한 뒤 역파싱 검증·근거 게이트·번호 정합 게이트를 통과한 문항만 남긴 학습 데이터. AMA-Bench 의 네 유형(A 회상 · B 인과 · C 상태 갱신 · D 상태 추상화)을 겨냥한다.
B 유형의 "왜"·"실패 뒤 다음 행동" 문항은 궤적에 기록된 에이전트의 이유(Thought) 를 근거로 한다.
프롬프트는 AMA compaction_v3_nostate 하네스 형식(Task / Step Index / Most Recent(Thought: 줄 포함) / Recalled / Questions /… See the full description on the dataset page: https://huggingface.co/datasets/HBKenerzai/agentic-state-recall-v1.dx3-recall-generalization-benchmark
Dx3 Recall Generalization Benchmark — held-out v0
Author: Asif Waliuddin · NXTG.AI · CC BY 4.0
A held-out recall goldenset (20 query→target pairs), disjoint from the tuning set, built to test whether a retrieval fix generalizes rather than memorizes the set it was tuned against. Every target was full-text-confirmed present in the live store before inclusion — so a recall@5 miss means weak retrieval/ranking, not absence. Honest methodology artifact: a generalization gate, not a… See the full description on the dataset page: https://huggingface.co/datasets/skinny-cloud/dx3-recall-generalization-benchmark.Nikolas-Memory-Recall-DatasetFIXEDfunc_recallfda-drug-recalls-2023-gpt
Structured FDA Drug Recalls — GPT-Enriched (2023)
This dataset contains structured FDA drug recall information for 2023. The data was extracted from FDA enforcement reports and enriched using GPT-4 to provide machine-readable formats in both CSV and JSON.
Included Fields:
Recall number
Classification (I, II, III, or Pending)
Product description
Manufacturer
Reason for recall (summarized)
Lot/expiration info
Distribution geography
Event initiation and enforcement dates… See the full description on the dataset page: https://huggingface.co/datasets/michaelrrosemba/fda-drug-recalls-2023-gpt.sera-4.5-django-t2-recall05-toolcalls
SERA-4.5A Django T2 (Recall=0.5) Toolcalls
This dataset contains normalized multi-turn tool-calling trajectories derived from:
Source dataset: allenai/Sera-4.5A-Django-T2
Filter: line_level_recall == 0.5
Splits
train.jsonl: 6200 records
val.jsonl: 331 records
Format
Each line is a JSON object with:
id: trajectory id
messages: normalized chat/tool-call messages
metadata: includes instance_id, func_name, func_path, line_level_recall
Processing… See the full description on the dataset page: https://huggingface.co/datasets/endsky/sera-4.5-django-t2-recall05-toolcalls.recall_intelligence_signals
Recall Intelligence Signals v1
Public recall and enforcement records prepared by SingleFoundry. Built for institutional_research buyers, this SingleFoundry data product packages 1 validated records with governed source evidence, quality checks, audit traceability, and ready-to-use delivery metadata.
Dataset Files
data/recall-intelligence-signals-extracted-records-csv.csv: latest validated SingleFoundry CSV package.
singlefoundry-metadata.json: release metadata… See the full description on the dataset page: https://huggingface.co/datasets/singlefoundry/recall_intelligence_signals.
