datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multisource-memory-benchmark
Multi-Source Memory Benchmark
Status — anonymous artefact for double-blind review (NeurIPS 2026 Evaluations & Datasets Track).
Author identities, organisations, and funders are intentionally withheld until the review period concludes.
A diagnostic testbed for selective question-answering (ANSWER / SKIP) over conflicting multi-source personal memory.
Each persona has five evidence streams projected from a single latent event table with known, controlled per-source distortions… See the full description on the dataset page: https://huggingface.co/datasets/anon-neuripsed26/multisource-memory-benchmark.Dans-MemoryCore-CoreCurriculum-Small
Dan's Memory Core: Core Curriculum Small
Broad strokes
This dataset aims to provide a foundation of knowledge common to a number of fields and areas of study. The question answer pairs were generated using a RAG implementation and a curated selection of source material. Ideally this will be the first in a series of datasets that will cover a wide range of topics.
Nomic Atlas Visualiztion
Cluster visualization for the dataset available here.
Topics… See the full description on the dataset page: https://huggingface.co/datasets/PocketDoc/Dans-MemoryCore-CoreCurriculum-Small.ISETrace-Memory-Queries
ISETrace Memory Query Corpus
This repository contains the frozen natural-language memory-query corpus used for exact-span retrieval experiments over ISETrace, an execution-grounded corpus of operating-system agent trajectories.
The release contains 6,597 English queries. Each query asks for information recoverable from one trajectory and identifies minimal answer-bearing quotes in stable source sections. Query records and authoring metadata are kept separate so metadata such as… See the full description on the dataset page: https://huggingface.co/datasets/HazeLocus/ISETrace-Memory-Queries.MemoryCraft
MemoryCraft — Unified Agent-Memory Benchmark Collection
Five memory benchmarks reformatted into one common schema for evaluating how
well an agent uses long-term memory. Two configs:
full/ — every instance of each source, unified.
selected/ — the evaluation subset used in our runs (QA balanced across
benchmarks; Membench = its largest/long-context instances).
benchmark
full instances
full QA
selected instances
selected QA
locomo
10
1986
10
1986
longmemeval
500
500… See the full description on the dataset page: https://huggingface.co/datasets/daven3/MemoryCraft.agent-memory-benchmark
Agent Memory Compression & Evaluation Benchmark
This dataset is a controlled evaluation testbed designed to benchmark long-term memory architectures for conversational AI agents. It stress-tests how agents handle long conversations with complex fact dynamics.
Dataset Structure
1. conversation.json
A 100-turn synthetic conversation (50 user, 50 assistant turns) containing embedded facts categorized under:
Simple Facts: Baseline retrieval details.… See the full description on the dataset page: https://huggingface.co/datasets/kushalicious/agent-memory-benchmark.
