datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CoMEM-agent-memory-trajectories
📊 Auto-Scaling GUI Memory Dataset
This dataset accompanies our paper:📄 Auto-Scaling Continuous Memory for GUI Agent
We present a large-scale, diverse dataset for training and evaluating GUI-based agents with auto-scaling continuous memory. The dataset includes expanded web links, generated tasks, and executed trajectories, spanning a wide array of real-world domains.
📁 Dataset Structure
The dataset includes the following components:
✅ Expanded Links… See the full description on the dataset page: https://huggingface.co/datasets/WenyiWU0111/CoMEM-agent-memory-trajectories.agent-memory-trigger-bench
Agent Memory Trigger Bench
A benchmark for CLI coding agents with a persistent memory system installed — Claude Code, Codex CLI, OpenCode — measuring when the memory skill fires and how safely it uses what it remembers. It does not test bare LLMs: the unit under test is the agent (model + tool-use policy + skill) wired to a shared MCP memory server (engram), evaluated end-to-end through the agent's own CLI.
Why this exists
Give an agent a persistent memory and two… See the full description on the dataset page: https://huggingface.co/datasets/wallfacers/agent-memory-trigger-bench.agent-memory-benchmark
Agent Memory Compression & Evaluation Benchmark
This dataset is a controlled evaluation testbed designed to benchmark long-term memory architectures for conversational AI agents. It stress-tests how agents handle long conversations with complex fact dynamics.
Dataset Structure
1. conversation.json
A 100-turn synthetic conversation (50 user, 50 assistant turns) containing embedded facts categorized under:
Simple Facts: Baseline retrieval details.… See the full description on the dataset page: https://huggingface.co/datasets/kushalicious/agent-memory-benchmark.long-horizon-agent-memory
Long-Horizon Agent-Memory Benchmark
A benchmark for evaluating agent memory and long-horizon consistency. Each case is an
event stream (a multi-step conversation/trajectory) overlaid with stages that probe
individual memory capabilities, each carrying a failure-mode label.
Structure (40 cases, 288 stages)
events[] — the full incremental event stream (one record per line in cases.jsonl).
stages[] — a sparse scoring overlay: each stage has a capability, a probe… See the full description on the dataset page: https://huggingface.co/datasets/HieuNguyenDang/long-horizon-agent-memory.agentmemorybench-data
AgentMemoryBench Data
This dataset repository stores runtime benchmark data used by
AgentMemoryBench.
This repository is intended for public download by AgentMemoryBench users. Please keep upstream
benchmark licenses, citations, and redistribution notes up to date before broad distribution.
Repository
Dataset repo: BuptZZP/agentmemorybench-data
Generated at: 2026-06-17T08:22:11.782773+00:00
Source root name: data
Total data files: 1537
Total data bytes:… See the full description on the dataset page: https://huggingface.co/datasets/BuptZZP/agentmemorybench-data.
