datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-memory-bench-corpus
agent-memory-bench: the experience corpus
The neutral feed for a preregistered, execution-graded benchmark of memory layers for coding
agents. Every memory product under test ingests these same bytes through its own write path,
then an agent is given real coding work in a real repository where success depends on something
established in an earlier session, and the artifact is graded by execution: the task's tests
pass or they do not.
There is no LLM judge anywhere in the primary… See the full description on the dataset page: https://huggingface.co/datasets/Gde05/agent-memory-bench-corpus.tool-reasoning-sft-MEMORY-mem_agent-sft-data-cleaned-rectified-408k
mem_agent-sft-data-cleaned-rectified
Multi-turn long-context memory-agent SFT dataset with explicit reasoning traces, structured tool calls, and sequential chunk-processing sub-chains.
Schema
Column
Type
Description
messages
string (JSON)
JSON-serialized list of {role, content} dicts. Roles: system, user, reasoning, tool_call, tool_output, answer
core_chain_OR_subcall
string
"core_chain" (full orchestration trace) or "subcall" (single chunk-processing step)… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-MEMORY-mem_agent-sft-data-cleaned-rectified-408k.mem_agent-model_based-rl-memoryagent-7b-ruler-qa-test-c27000-t1024-10s-agnosticrepro-learning-to-share-selective-memory-for-efficient-parallel-agentic-systems-traces
Agent traces
Agent sessions published from a Trackio Logbook.
mem_agent-model_based-rl-memoryagent-7b-ruler-qa-test-c2048-t1024-10s-agnostic-nocontextadaption-agent-memory-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-agent-memory (augmented)
This dataset contains samples of conversations between a user and an assistant, paired with the corresponding structured memory updates extracted for long-term storage. Each sample includes the full dialogue context, existing memory state, and the resulting JSON output containing new narrative summaries and atomic facts with specific keys and values.… See the full description on the dataset page: https://huggingface.co/datasets/huyxdang/adaption-agent-memory-augmented.agent-memory-benchmark
Agent Memory Compression & Evaluation Benchmark
This dataset is a controlled evaluation testbed designed to benchmark long-term memory architectures for conversational AI agents. It stress-tests how agents handle long conversations with complex fact dynamics.
Dataset Structure
1. conversation.json
A 100-turn synthetic conversation (50 user, 50 assistant turns) containing embedded facts categorized under:
Simple Facts: Baseline retrieval details.… See the full description on the dataset page: https://huggingface.co/datasets/kushalicious/agent-memory-benchmark.adaption-agent-memory-augmented-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-agent-memory (augmented)
This dataset contains samples of conversations between a user and an assistant, paired with the corresponding structured memory updates extracted for long-term storage. Each sample includes the full dialogue context, existing memory state, and the resulting JSON output containing new narrative summaries and atomic facts with specific keys and values.… See the full description on the dataset page: https://huggingface.co/datasets/huyxdang/adaption-agent-memory-augmented-v1.long-horizon-agent-memory
Long-Horizon Agent-Memory Benchmark
A benchmark for evaluating agent memory and long-horizon consistency. Each case is an
event stream (a multi-step conversation/trajectory) overlaid with stages that probe
individual memory capabilities, each carrying a failure-mode label.
Structure (40 cases, 288 stages)
events[] — the full incremental event stream (one record per line in cases.jsonl).
stages[] — a sparse scoring overlay: each stage has a capability, a probe… See the full description on the dataset page: https://huggingface.co/datasets/HieuNguyenDang/long-horizon-agent-memory.agent-memory-graph
Agent Memory Graph
Unified knowledge graph combining structured entity data from mcp-agents-ark and semantic content from mempalace (ChromaDB vector store). Sanitized for public release.
Usage
from datasets import load_dataset
# Unified (27 cols) — use this for most training tasks
ds = load_dataset("Ev3lynx727/agent-memory-graph", "unified", split="train")
# Ark entities only (12 cols) — structured entity graph
ds_ark =… See the full description on the dataset page: https://huggingface.co/datasets/Ev3lynx727/agent-memory-graph.ai-agent-memory
AI Agent Memory Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP
This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc.
The dataset is helpful for AI… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/ai-agent-memory.agent-memory-research-corpus
Agent Memory Research Corpus (AMRC)
A public, citable dataset for agent memory research and systems.
This dataset catalogues papers, systems, benchmarks, and design patterns related to long-term memory in autonomous agents. It is intended to serve as a canonical reference corpus for researchers and practitioners building memory-augmented agents.
Dataset Summary
Field
Value
Repository
https://huggingface.co/datasets/trentdoney/agent-memory-research-corpus… See the full description on the dataset page: https://huggingface.co/datasets/trentdoney/agent-memory-research-corpus.mem_agent-bertscore-rl-memoryagent-14b-docfinqa-train-c4096-t4096-1000s-agnosticmemory-agent-sft-v1mem_agent-model_based-rl-memoryagent-14b-infbench-longbook-qa-test-c31000-t4096-1000s-agnosticmem_agent-model_based-rl-memoryagent-7b-ruler-qa-test-c2048-t1024-10s-agnostic-fullcontextmem_agent-model_based-rl-memoryagent-7b-ruler-qa-test-c2048-t1024-10s-agnosticmem_agent-model_based-rl-memoryagent-7b-infbench-code-debug-test-c8192-t4096-1000s-a-fullcontextemgena_agent_vector_memory_leak_guard_mcp_teaser
🚀 Emgena Agent Vector Memory & GraphRAG Desync MCP (Turnkey MCP Plugin & Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Scenarios)🏆 Get the Full Turnkey MCP Plugin & Production Master Package (1,000 Samples) on Gumroad:👉 Purchase Turnkey MCP Plugin on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
🌟 Plugin Overview & Features
Turnkey Semantic Drift Quarantine, Episodic Memory Decay Checker & Vector Hallucination Guard for Cursor &… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_agent_vector_memory_leak_guard_mcp_teaser.mem_agent-model_based-rl-memoryagent-14b-triviaqa-llama-memorization-val-c4096-t2048-1000s-agnosmem_agent-model_based-rl-memoryagent-7b-barexamqa-train-c32000-t128-1000s-agnosticmem_agent-model_based-rl-memoryagent-14b-triviaqa-llama-memorization-val-c27000-t2048-1000s-agnomem_agent-bertscore-rl-memoryagent-14b-housingqa-test-c256-t256-20s-agnosticmem_agent-model_based-rl-memoryagent-7b-infbench-code-debug-test-c8192-t4096-1000s-agnosticmem_agent-model_based-rl-memoryagent-7b-infbench-code-debug-test-c8192-t4096-10s-agn-nocontextstrl-main-ec-clawbench_flawedflowed-gc-claude_query-mc-claude_agent_sonnet-rbc-memory_bm-r0mem_agent-bertscore-rl-memoryagent-14b-housingqa-test-c512-t256-1000s-agnosticmem_agent-bertscore-rl-memoryagent-14b-barexamqa-train-c256-t128-20s-agnosticmem_agent-bertscore-rl-memoryagent-14b-gsminf-ops_6-c4096-t2048-20s-agnosticmem_agent-model_based-rl-memoryagent-7b-barexamqa-train-c256-t128-1000s-agnostic-nocontext
