replication
replicationbenchfactprobe-replication-negatives-allnames-v1
Plausible wrong answers, asked under every name (42,267,800 rows)
Status: final — 40 of 40 runs. Models present:
13b, 7b. Training stages present: s1, s2, s3, s4, s5.
What this fixes
A real fact is put to the model under the full cross product of the two
people's name lists, and counts as recognised if any one combination gets a
Yes. That is He et al.'s rule. Until 2026-08-26 the wrong answer it was
compared against was asked under one name per person, so the… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-negatives-allnames-v1.ReplicationBench
ReplicationBench
arXiv: ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers?
GitHub: https://github.com/Christine8888/replicationbench-release
Dataset Description
The ReplicationBench dataset contains 111 astrophysics research replication tasks, spanning complete replications of 20 research papers. The dataset includes:
Original and masked manuscript text
Metadata (title, abstract, publication info, etc.)
Pointers to datasets and dataset access… See the full description on the dataset page: https://huggingface.co/datasets/ChristineYe8/ReplicationBench.lifemem-replication-src
LifeMem replication (arXiv:2608.19621)
Replication of Mitigating Identity Essentialism in LLM Agents with Longitudinal Life Trajectories (Wang, Zhou, Du, Su, Cao, Pan, Ai, Wu, Zhang, Liu). Official code: halsayxi/LifeMem.
The paper's claim: static demographic prompts make LLM survey agents essentialist — within-group answers collapse, SES clusters separate. LifeMem stores life events in a hippocampal retriever (top-K=5, α=0.9) and consolidates them into a per-agent LoRA adapter.… See the full description on the dataset page: https://huggingface.co/datasets/mtorres98/lifemem-replication-src.crimsonred-paper-replication
CrimsonRed — Cross-Architecture Emotion-Prime Steering Replication
Replication of the emotion-prime steering protocol from arXiv:2607.18691 (NSM
semantic primes as explanans for emotion in LLMs), extended across four
architectures. Generated by scripts/paper_faithful_steering.py in the
CrimsonRed project.
The finding
The paper's core claim — that semantic-prime recipe directions steer emotion more
strongly than Scherer appraisal directions — replicates… See the full description on the dataset page: https://huggingface.co/datasets/musicakamusic/crimsonred-paper-replication.gz3d-aion-replication-attempt
GalaxyZoo 3D Segmentation Benchmark
An attempted replication of the GZ3D segmentation dataset used in section 7.2.3 in the AION-1. This dataset contains volunteer segmentations for the following channels: center, star, spiral, bar. It also contains the RGB images from Legacy Survey and the tokens from the AION-1 image tokenizer of the Legacy Survey image bands. Notably, these are pre–AION-1 transformer encoder, but post image-codec encoder/quantization. The reason why it's an… See the full description on the dataset page: https://huggingface.co/datasets/astronolan/gz3d-aion-replication-attempt.
