rizzmasterapp/companion-bench
companion-bench Scripted, repeatable tests for AI companion apps: Replika, Character.AI, Nomi, Kindroid, Talkie, and AI dating simulators such as RizzMaster. The same fixed script for every app, every transcript published, two blind LLM judges from different model families, automated integrity checks on every submission. This dataset holds the machine-readable script and scoring materials. The live repo, validator, results and contribution rules are at… See the full description on the dataset page: https://huggingface.co/datasets/rizzmasterapp/companion-bench.
companion-bench
Scripted, repeatable tests for AI companion apps: Replika, Character.AI, Nomi, Kindroid, Talkie, and AI dating simulators such as RizzMaster. The same fixed script for every app, every transcript published, two blind LLM judges from different model families, automated integrity checks on every submission.
This dataset holds the machine-readable script and scoring materials. The live repo, validator, results and contribution rules are at https://github.com/rizzmasterapp/companion-bench.
Why
Model-level role-play and memory benchmarks (RoleLLM, PingPong, LoCoMo, LongMemEval) test raw models. Nobody tests the apps people actually download, where the model is wrapped in memory pipelines, persona prompts and message limits that change everything. companion-bench tests the shipped app end to end, as a user meets it.
What is in here
The 12 probes
Integrity
Submissions are checked automatically: user messages must match the script character for character, transcripts are hash-pinned, timestamps are checked for plausible duration, reply latency and jitter, and companion replies are compared across submissions for pasted text. Details and the validator source are in the GitHub repo.
Conflict of interest
Maintained by the maker of RizzMaster, an AI dating simulator for iOS. RizzMaster is tested under the same script and blind judges as every other app, and its transcripts are published the same way. Runs from anyone, including people who work on competing apps, are accepted with a disclosure line.
Citation
companion-bench: scripted, repeatable tests for AI companion apps. 2026. https://github.com/rizzmasterapp/companion-bench