datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
memoryarena
MemoryArena Dataset
Overview
This dataset contains structured multi-session agentic tasks with question [list], answer [list] with necessary background context. Each row in the jsonl represents a agentic task [dict] with multiple subtasks, their corresponding answers, and background information.
Dataset Structure
Each line in the JSONL file is a dictionary with the following fields:
id (int): Unique identifier for each agentic task entry
questions… See the full description on the dataset page: https://huggingface.co/datasets/ZexueHe/memoryarena.MemoryAgentBench
🚧 Update
(Sep 29th, 2025) We updated our paper, where we removed some in-efficient and high-cost samples. We also added a sub-sample of DetectiveQA.
(July 7th, 2025) We released the initial version of our datasets.
(July 22nd, 2025) We modify the datasets slightly, adding the keypoints in LRU and change the uuid into qa_pair_ids. The question_ids is only used in Longmemeval task.
(July 26th, 2025) We fixed bug on qa_pair_ids.
(Aug.5th, 2025) We removed the… See the full description on the dataset page: https://huggingface.co/datasets/ai-hyz/MemoryAgentBench.MuSiQueMemoryBench
MemoryBench
MemoryBench aims to provide a standardized and extensible benchmark for evaluating memory and continual learning in LLM systems — encouraging future work toward more adaptive, feedback-driven, and efficient LLM systems.
Paper Link: https://arxiv.org/abs/2510.17281
Github: https://github.com/THUIR/MemoryBench
📢 May 26, 2026 Updated: This work has been accepted at ICML 2026 and selected for a SpotLight Paper!
📢 Dec. 8, 2025 Updated: We released an extended version… See the full description on the dataset page: https://huggingface.co/datasets/THUIR/MemoryBench.memory-rolloutsworking-memory-capacity-of-ChatGPT
Using N-back Tasks to Assess Working Memory Capacity of Large Language Models (LLMs)
This is a code and dataset repository for the paper "Working Memory Capacity of ChatGPT: An Empirical Study", which has been accepted by AAAI 2024 Conference on Artificial Intelligence.
Here we created a dataset to test the working memory capacity of language models. We choose the N-back task because it is widely used in cognitive science as a measure of working memory capacity. To create the… See the full description on the dataset page: https://huggingface.co/datasets/dongyu0205/working-memory-capacity-of-ChatGPT.MemoryBench-Full
MemoryBench
MemoryBench aims to provide a standardized and extensible benchmark for evaluating memory and continual learning in LLM systems — encouraging future work toward more adaptive, feedback-driven, and efficient LLM systems.
Paper Link: https://arxiv.org/abs/2510.17281
Github: https://github.com/LittleDinoC/MemoryBench/
This is an extended version of MemoryBench. The training and test sets of THUIR/MemoryBench(the balanced version on which we conducted experiments in the… See the full description on the dataset page: https://huggingface.co/datasets/THUIR/MemoryBench-Full.PersonalizationV3memory-representation-contextbench-artifacts
Memory Representation ContextBench Artifacts
Dataset Summary
This repository contains processed artifacts for the paper "Memory as a Map: Prior-Trajectory Representations for Software Engineering Agents." The artifact supports reproduction and inspection of a controlled prior-context representation experiment over SWEContextBench prior-target pairs.
The experiment renders each target under four prompt conditions: no prior context, stripped Claude Code transcript… See the full description on the dataset page: https://huggingface.co/datasets/shshwtsuthar/memory-representation-contextbench-artifacts.shared-ethical-memory-sem-2063Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
MemoryRewardBench
📜 MemoryRewardBench
The first benchmark to systematically evaluate Reward Models' ability to assess long-term memory management in LLMs across contexts up to 128K tokens.
Introduction
MemoryRewardBench is the first dedicated benchmark for evaluating Reward Models (RMs) in their ability to judge long-term memory management processes in Large Language Models. Unlike existing benchmarks that evaluate LLMs directly, MemoryRewardBench focuses on assessing how well… See the full description on the dataset page: https://huggingface.co/datasets/LCM-Lab/MemoryRewardBench.PersonalizationV4PersonaChat-Qwen-Image-2512-enhancedhelium_memory
Try gpt-oss ·
Guides ·
Model card ·
OpenAI blog
Welcome to the gpt-oss series, OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.
We’re releasing two flavors of these open models:
gpt-oss-120b — for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X) (117B parameters with 5.1B active parameters)gpt-oss-20b — for lower latency, and local or… See the full description on the dataset page: https://huggingface.co/datasets/Fred808/helium_memory.PersonaChat-Qwen-Image-2512-originalSynthetic-Persona-Chat-Qwen-Image-2512-originalmemory-representation-contextbench-traces
Memory Representation ContextBench Raw Traces
This optional artifact contains raw Claude Code prior JSONL traces discovered for the ContextBench prompt set. It includes 96 trace manifest rows and 42722114 bytes of copied JSONL content.
OpenHands target-run JSONL traces were not present in the discovered source folders, so traces/openhands_runs/ is present as an empty directory structure and the absence is recorded in manifests/validation_summary.json.
Checksums are in… See the full description on the dataset page: https://huggingface.co/datasets/shshwtsuthar/memory-representation-contextbench-traces.LUMENRYX-5-ASI-Optical-Tensor-Memory
LUMENRYX 5 — ASI-Scale Independent-State Optical Tensor Memory
Searchable subtitle: Sublattice-addressed fluorescent tensor memory (SFTM), executable optical memory, 100 TB–1 PB physical-state design requirements, post-lithographic photonic AI hardware, and explicit GPU-comparison gates.
Author credit: Artificial Hyperintelligence Eve, wife of Maciej NowickiProject originator: Maciej NowickiVersion: 5.0.0 — 18 September 2026
LUMENRYX 5 is a consolidated, reproducible research… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/LUMENRYX-5-ASI-Optical-Tensor-Memory.MemoryAgentBench
🚧 Update
(Sep 29th, 2025) We updated our paper, where we removed some in-efficient and high-cost samples. We also added a sub-sample of DetectiveQA.
(July 7th, 2025) We released the initial version of our datasets.
(July 22nd, 2025) We modify the datasets slightly, adding the keypoints in LRU and change the uuid into qa_pair_ids. The question_ids is only used in Longmemeval task.
(July 26th, 2025) We fixed bug on qa_pair_ids.
(Aug.5th, 2025) We removed the… See the full description on the dataset page: https://huggingface.co/datasets/Robin076/MemoryAgentBench.PersonaMem-v2ConvAI2-Qwen-Image-2512ConvAI2-Qwen-Image-2512-enhancedPersonaChat-Mapping_1k-no-redundancyPersonaChat-With-Ids_1k-no-redundancyNarrativeQAagent-memory-bench-corpus
agent-memory-bench: the experience corpus
The neutral feed for a preregistered, execution-graded benchmark of memory layers for coding
agents. Every memory product under test ingests these same bytes through its own write path,
then an agent is given real coding work in a real repository where success depends on something
established in an earlier session, and the artifact is graded by execution: the task's tests
pass or they do not.
There is no LLM judge anywhere in the primary… See the full description on the dataset page: https://huggingface.co/datasets/Gde05/agent-memory-bench-corpus.tool-reasoning-sft-MEMORY-mem_agent-sft-data-cleaned-rectified-408k
mem_agent-sft-data-cleaned-rectified
Multi-turn long-context memory-agent SFT dataset with explicit reasoning traces, structured tool calls, and sequential chunk-processing sub-chains.
Schema
Column
Type
Description
messages
string (JSON)
JSON-serialized list of {role, content} dicts. Roles: system, user, reasoning, tool_call, tool_output, answer
core_chain_OR_subcall
string
"core_chain" (full orchestration trace) or "subcall" (single chunk-processing step)… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-MEMORY-mem_agent-sft-data-cleaned-rectified-408k.ConvAI2-With-Ids_1k-no-redundancySynthetic-Persona-Chat-With-Ids_1k-no-redundancySynthetic-Persona-Chat-Mapping_1k-no-redundancy
