Theory_of_mind
Tiny_Theory_of_Mind
Tiny Theory of Mind
Tiny Theory of Mind is our first attempt at evaluating theory of mind capabilities in small language models. The benchmark covers a wide variety of ToM topics, ranging in difficulties that, for humans, would be appropriate for Pre-K through 6th grade.
The benchmark is designed primarily for base-model continuation log-likelihood scoring. It does not require instruction following, chain-of-thought, or generated explanations. Random-choice accuracy is 25%.… See the full description on the dataset page: https://huggingface.co/datasets/AxiomicLabs/Tiny_Theory_of_Mind.Theory_of_Mind_CoMMET
CoMMET
CoMMET is a multi-turn, multimodal benchmark designed to evaluate the Theory of Mind (ToM) capabilities of multimodal large language models (MLLMs).
Unlike conventional single-turn Theory of Mind benchmarks, CoMMET represents each scenario as a sequence of interconnected turns. Models are required to reason about stories, previous interactions, feedback, questions, and, when necessary, visual information.
The benchmark covers multiple types of mental-state reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/rrchen2026/Theory_of_Mind_CoMMET.theory-of-mindQ&A testing theory of mind, in Alpaca format, generated by gpt-4-1106-preview. OpenAI terms apply.
Each answer was double-checked by gpt-4-1106-preview, and suspicious answers were removed, since even GPT4 struggles with accuracy in this test. This does not guarantee that the remaining entries are correct, but the accuracy should be better than base.
Files:
theory_of_mind.json Usual, double-checked TOM Q&A (150 entries)
theory_of_mind_longer.json Slightly longer Q&A (50 entries)… See the full description on the dataset page: https://huggingface.co/datasets/grimulkan/theory-of-mind.grimulkan_theory-of-mind-ShareGPTTheory-of-mind-reasoning-packThis dataset is Neurvance.com Theory of mind reasoning pack. By downloading you agree to Neurvance Policies https://neurvance.com/policy.html For a compliance pack for AI Article 10 EU regulations, visit https://neurvance.com/contact.html
theory-of-mind-dpoThis is grimulkan/theory-of-mind with "rejected" responses generated using mistralai/Mistral-7B-Instruct-v0.2, and the file formatted for use in DPO training.
The code used to generate the dataset can be found in this repository: https://github.com/DocShotgun/LLM-datagen
