datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MemoryAgentBench
🚧 Update
(Sep 29th, 2025) We updated our paper, where we removed some in-efficient and high-cost samples. We also added a sub-sample of DetectiveQA.
(July 7th, 2025) We released the initial version of our datasets.
(July 22nd, 2025) We modify the datasets slightly, adding the keypoints in LRU and change the uuid into qa_pair_ids. The question_ids is only used in Longmemeval task.
(July 26th, 2025) We fixed bug on qa_pair_ids.
(Aug.5th, 2025) We removed the… See the full description on the dataset page: https://huggingface.co/datasets/ai-hyz/MemoryAgentBench.MuSiQuelldms-associative-memory-samples
LLDMs Associative Memory — Generated Samples
Model-generated text for the paper:
Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data
Bao Pham, Mohammed J. Zaki, Luca Ambrogioni, Dmitry Krotov, Matteo Negri
Accepted to EMNLP 2026 (Main Conference).
arXiv:2604.26841 · paper · code · checkpoints
29.5 million generated sequences (~3.8B tokens) sampled from the released checkpoints — one
generation run per (model size, training-set fraction). These… See the full description on the dataset page: https://huggingface.co/datasets/lemoncmd/lldms-associative-memory-samples.memory-rolloutsPersonalizationV3memory-representation-contextbench-artifacts
Memory Representation ContextBench Artifacts
Dataset Summary
This repository contains processed artifacts for the paper "Memory as a Map: Prior-Trajectory Representations for Software Engineering Agents." The artifact supports reproduction and inspection of a controlled prior-context representation experiment over SWEContextBench prior-target pairs.
The experiment renders each target under four prompt conditions: no prior context, stripped Claude Code transcript… See the full description on the dataset page: https://huggingface.co/datasets/shshwtsuthar/memory-representation-contextbench-artifacts.PersonalizationV4memory_layers
CorpusQA-Films — aggregation-QA dataset (v1)
Synthetic corpus-level aggregation questions over English Wikipedia film articles, with
self-distilled chain-of-thought. Built to train the memory-layers model (frozen Qwen3-4B +
learnable retrieval/memory layer). Formatted to match
ragrawal36/multihop_qa_sft-hard-neg-cot.
Files (HF-ready)
file
schema
rows
corpusqa_films_qa.parquet
question:str, answer:str, pos_doc_ids:list<int32>, neg_doc_ids:list<int32>… See the full description on the dataset page: https://huggingface.co/datasets/jordanlin/memory_layers.PersonaChat-Qwen-Image-2512-enhancedPersonaChat-Qwen-Image-2512-originalhelium_memory
Try gpt-oss ·
Guides ·
Model card ·
OpenAI blog
Welcome to the gpt-oss series, OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.
We’re releasing two flavors of these open models:
gpt-oss-120b — for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X) (117B parameters with 5.1B active parameters)gpt-oss-20b — for lower latency, and local or… See the full description on the dataset page: https://huggingface.co/datasets/Fred808/helium_memory.Synthetic-Persona-Chat-Qwen-Image-2512-originalmemory-representation-contextbench-traces
Memory Representation ContextBench Raw Traces
This optional artifact contains raw Claude Code prior JSONL traces discovered for the ContextBench prompt set. It includes 96 trace manifest rows and 42722114 bytes of copied JSONL content.
OpenHands target-run JSONL traces were not present in the discovered source folders, so traces/openhands_runs/ is present as an empty directory structure and the absence is recorded in manifests/validation_summary.json.
Checksums are in… See the full description on the dataset page: https://huggingface.co/datasets/shshwtsuthar/memory-representation-contextbench-traces.MemoryAgentBench
🚧 Update
(Sep 29th, 2025) We updated our paper, where we removed some in-efficient and high-cost samples. We also added a sub-sample of DetectiveQA.
(July 7th, 2025) We released the initial version of our datasets.
(July 22nd, 2025) We modify the datasets slightly, adding the keypoints in LRU and change the uuid into qa_pair_ids. The question_ids is only used in Longmemeval task.
(July 26th, 2025) We fixed bug on qa_pair_ids.
(Aug.5th, 2025) We removed the… See the full description on the dataset page: https://huggingface.co/datasets/Robin076/MemoryAgentBench.ConvAI2-Qwen-Image-2512PersonaMem-v2MemoryMatters_DatasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 50,
"total_frames": 14325,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/wayrise/MemoryMatters_Dataset.ConvAI2-Qwen-Image-2512-enhancedPersonaChat-Mapping_1k-no-redundancyPersonaChat-With-Ids_1k-no-redundancyNarrativeQAmemory-chattool-reasoning-sft-MEMORY-mem_agent-sft-data-cleaned-rectified-408k
mem_agent-sft-data-cleaned-rectified
Multi-turn long-context memory-agent SFT dataset with explicit reasoning traces, structured tool calls, and sequential chunk-processing sub-chains.
Schema
Column
Type
Description
messages
string (JSON)
JSON-serialized list of {role, content} dicts. Roles: system, user, reasoning, tool_call, tool_output, answer
core_chain_OR_subcall
string
"core_chain" (full orchestration trace) or "subcall" (single chunk-processing step)… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-MEMORY-mem_agent-sft-data-cleaned-rectified-408k.ConvAI2-With-Ids_1k-no-redundancySynthetic-Persona-Chat-With-Ids_1k-no-redundancynyush-galaxea-a1-kitchen-no-memory
NYUSH Galaxea A1 — Kitchen
Formats and branches
Branch
Representation
main (default)
EEF · LeRobot v2.1
v3
EEF · LeRobot v3.0
Preview
Each GIF contains both camera views on one shared playback clock. Left: Agent View; right: Wrist Camera. The timestamp shows elapsed dataset time; playback is 2× speed.
Put the pot on the stove and turn on the switch
Episode 20 · complete episode · 24.33 s recorded.
Pour… See the full description on the dataset page: https://huggingface.co/datasets/pengyue-polaron/nyush-galaxea-a1-kitchen-no-memory.Synthetic-Persona-Chat-Mapping_1k-no-redundancyPersonaChat-Mapping_1k-Qwen-original-no-redundancySynthetic-Persona-Chat-Mapping_1k-ERNIE-originalConvAI2-Mapping_1k-no-redundancy
