datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MemoryBench
MemoryBench
MemoryBench aims to provide a standardized and extensible benchmark for evaluating memory and continual learning in LLM systems — encouraging future work toward more adaptive, feedback-driven, and efficient LLM systems.
Paper Link: https://arxiv.org/abs/2510.17281
Github: https://github.com/THUIR/MemoryBench
📢 May 26, 2026 Updated: This work has been accepted at ICML 2026 and selected for a SpotLight Paper!
📢 Dec. 8, 2025 Updated: We released an extended version… See the full description on the dataset page: https://huggingface.co/datasets/THUIR/MemoryBench.MemoryBench-Full
MemoryBench
MemoryBench aims to provide a standardized and extensible benchmark for evaluating memory and continual learning in LLM systems — encouraging future work toward more adaptive, feedback-driven, and efficient LLM systems.
Paper Link: https://arxiv.org/abs/2510.17281
Github: https://github.com/LittleDinoC/MemoryBench/
This is an extended version of MemoryBench. The training and test sets of THUIR/MemoryBench(the balanced version on which we conducted experiments in the… See the full description on the dataset page: https://huggingface.co/datasets/THUIR/MemoryBench-Full.memorybench
MemoryBench Dataset
MemoryBench is a benchmark dataset designed to evaluate spatial memory and action recall in robotic manipulation. This dataset accompanies the SAM2Act+ framework, introduced in the paper SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation. For detailed task descriptions and more information about this paper, please visit SAM2Act's website. Code can be found at https://github.com/sam2act/sam2act.
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/hqfang/memorybench.MemoryBench-ResultsMemoryBench Experiment Results
Paper •
Code •
Dataset
Overview
This repository contains experiment results for
MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems. MemoryBench evaluates whether LLM systems can learn from accumulated user
feedback during service time. The official benchmark data is hosted at
THUIR/MemoryBench.
This repository is an artifact archive for published runs.
It stores model predictions, per-sample evaluation details… See the full description on the dataset page: https://huggingface.co/datasets/THUIR/MemoryBench-Results.icml2026-64918-memorybench-repro-resultsMemoryBench-Fullmemory-bench
memory-bench v1.0.0, public release
A screened benchmark for organizational memory in agent harnesses: does a
memory system keep a rule that was stated once, drop a fact that was
superseded, and pick the right one when tiers conflict?
371 valid paired probe instances across 3 simulated
organizations, drawn from 486 probes over 612 events. Scored as pair
credit: an instance counts only if the base task and its counterfactual twin
both pass, so anything answerable from priors… See the full description on the dataset page: https://huggingface.co/datasets/notmehul/memory-bench.memorybench-lerobotmemorybench-extended
MemoryBench Extended
Extension of MemoryBench with additional memory-dependent manipulation tasks. The original three MemoryBench tasks (put_block_back, reopen_drawer, rearrange_block) are not duplicated here — download those from hqfang/memorybench.
Task: stack_and_swap
Two identically colored blocks (A, B) start on two colored patches. The robot must:
Pick block A and place it at the center.
Pick block B and stack it on top of A at the center.
Press a button (memory… See the full description on the dataset page: https://huggingface.co/datasets/phanikiran1169/memorybench-extended.MemoryBench-off-policy-results
