Agent Memory
agent-memory-compaction-trajectories
Agent Memory Compaction Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/agent-memory-compaction-trajectories.CoMEM-agent-memory-trajectories
📊 Auto-Scaling GUI Memory Dataset
This dataset accompanies our paper:📄 Auto-Scaling Continuous Memory for GUI Agent
We present a large-scale, diverse dataset for training and evaluating GUI-based agents with auto-scaling continuous memory. The dataset includes expanded web links, generated tasks, and executed trajectories, spanning a wide array of real-world domains.
📁 Dataset Structure
The dataset includes the following components:
✅ Expanded Links… See the full description on the dataset page: https://huggingface.co/datasets/WenyiWU0111/CoMEM-agent-memory-trajectories.agent-memory-bench-corpus
agent-memory-bench: the experience corpus
The neutral feed for a preregistered, execution-graded benchmark of memory layers for coding
agents. Every memory product under test ingests these same bytes through its own write path,
then an agent is given real coding work in a real repository where success depends on something
established in an earlier session, and the artifact is graded by execution: the task's tests
pass or they do not.
There is no LLM judge anywhere in the primary… See the full description on the dataset page: https://huggingface.co/datasets/Gde05/agent-memory-bench-corpus.tool-reasoning-sft-MEMORY-mem_agent-sft-data-cleaned-rectified-408k
mem_agent-sft-data-cleaned-rectified
Multi-turn long-context memory-agent SFT dataset with explicit reasoning traces, structured tool calls, and sequential chunk-processing sub-chains.
Schema
Column
Type
Description
messages
string (JSON)
JSON-serialized list of {role, content} dicts. Roles: system, user, reasoning, tool_call, tool_output, answer
core_chain_OR_subcall
string
"core_chain" (full orchestration trace) or "subcall" (single chunk-processing step)… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-MEMORY-mem_agent-sft-data-cleaned-rectified-408k.daily-paper-2026-07-30-agent-memory-tiering-recall-cost
Memory Tiering Policies for Long-Running Autonomous Agent Harnesses: Mapping the Recall-Cost Frontier
TL;DR — On a production agent-memory corpus, semantic deduplication (token-Jaccard merging before ranking) outperforms both recency and frequency ordering by +17.4% AUC, reaching full recall coverage at 70% of the corpus budget. The shipped per-section item-count cap, not the character cap, is the binding constraint.
ThakiCloud AI Research · 2026-07-30 · 📝 Tech blog (KO)… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-30-agent-memory-tiering-recall-cost.mem_agent-model_based-rl-memoryagent-7b-ruler-qa-test-c27000-t1024-10s-agnostic
