datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MEME
MEME: Multi-Entity and Evolving Memory Evaluation
A benchmark for evaluating LLM memory systems along two orthogonal dimensions: entity scope (single vs. multi-entity) and temporal dynamics (static vs. evolving). MEME defines six tasks targeting memory-intensive operations in each quadrant, including two task types that no prior benchmark covers: Cascade (propagating updates through dependency rules) and Absence (recognizing uncertainty when a previously valid answer becomes… See the full description on the dataset page: https://huggingface.co/datasets/meme-benchmark/MEME.meta-meme
Meta-Meme Consultation URLs Dataset
Description
This dataset contains 2177 consultation URLs generated from the Meta-Meme formally verified system. Each URL represents a consultation with one of 9 AI muses about a specific file in the repository.
Dataset Structure
file: Path to the file in the repository
muse: Assigned AI muse (Calliope, Clio, Erato, Euterpe, Melpomene, Polyhymnia, Terpsichore, Thalia, Urania)
tool: Consultation tool (llm, lean4, rustc… See the full description on the dataset page: https://huggingface.co/datasets/introspector/meta-meme.MEME-fillers
MEME Benchmark — Filler Sessions
Filtered filler sessions used by the MEME memory benchmark for haystack assembly. Two domain-matched pools, both produced by length-filtering and LLM-judge conflict filtering against MEME's evidence entities.
Files
File
Domain
Sessions
Source
fillers_pl.json
Personal Life
1,009
LongMemEval-S haystack (non-evidence sessions, deduplicated)
fillers_sw.json
Software Project
9,008
ShareGPT 52K (English coding subset)… See the full description on the dataset page: https://huggingface.co/datasets/meme-benchmark/MEME-fillers.
