tombench
Datasets
All datasets matching “tombench”tombench-en
ToMBench (English-only, OLMES-ready mirror)
This is a clean, English-only mirror of ToMBench (Chen et al., ACL 2024).
Source: https://github.com/zhchen18/ToMBench Paper: arXiv:2402.15052
Why this mirror exists
The official ToMBench distribution is JSONL on GitHub with bilingual (Chinese/English)
fields and inconsistent type inference (some rows have option fields as numeric, others
as string — breaks datasets.load_dataset("json", ...)). This mirror:
Drops all… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/tombench-en.tom-benchmark
UniToMBench Dataset
Dataset Summary
UniToMBench is a unified benchmark designed to evaluate the Theory of Mind (ToM) reasoning abilities of large language models (LLMs). It integrates and extends existing ToM benchmarks by providing narrative-based multiple-choice questions (MCQs) that span a wide range of ToM tasks—including false belief reasoning, perspective taking, emotion attribution, scalar implicature, and more.
This dataset supports the research paper… See the full description on the dataset page: https://huggingface.co/datasets/Shamant23/tom-benchmark.tombench_merged
TomBench Merged Dataset (Exact Matching)
This dataset contains the merged results of TomBench evaluation with the original TomBench dataset, using exact string matching.
Dataset Statistics
Total records: 2860
Exact matches: 2860
Manual matches: 0
Average model score: 0.5066
Matching Strategy
This version uses exact string matching after text normalization:
Remove extra whitespace and normalize formatting
Match stories exactly between datasets
Report any… See the full description on the dataset page: https://huggingface.co/datasets/ycfNTU/tombench_merged.tombench_mergetombench-embeddingToMBench_Hard
