long-horizon-memory
long-horizon-agent-memory
Long-Horizon Agent-Memory Benchmark
A benchmark for evaluating agent memory and long-horizon consistency. Each case is an
event stream (a multi-step conversation/trajectory) overlaid with stages that probe
individual memory capabilities, each carrying a failure-mode label.
Structure (40 cases, 288 stages)
events[] — the full incremental event stream (one record per line in cases.jsonl).
stages[] — a sparse scoring overlay: each stage has a capability, a probe… See the full description on the dataset page: https://huggingface.co/datasets/HieuNguyenDang/long-horizon-agent-memory.humanoid-long-horizon-intent-memory
Long-Horizon Intent Memory
Captures how humanoids retain and adapt to long-term human goals.
