social
Datasets
All datasets matching “social”social_i_qaWe introduce Social IQa: Social Interaction QA, a new question-answering benchmark for testing social commonsense intelligence. Contrary to many prior benchmarks that focus on physical or taxonomic knowledge, Social IQa focuses on reasoning about people’s actions and their social implications. For example, given an action like "Jesse saw a concert" and a question like "Why did Jesse do this?", humans can easily infer that Jesse wanted "to see their favorite performer" or "to enjoy the music", and not "to see what's happening inside" or "to see if it works". The actions in Social IQa span a wide variety of social situations, and answer candidates contain both human-curated answers and adversarially-filtered machine-generated candidates. Social IQa contains over 37,000 QA pairs for evaluating models’ abilities to reason about the social implications of everyday events and situations. (Less)social-sim-bench-gensworldcup2026
⚽ WorldCup Arena
A Leakage-Free Forecasting Benchmark on a Live Tournament
Can a language model forecast a match — when the match had not been played at the moment it was asked?
🌐 Language / 语言 : 中文 ▾
📊 四张表
点开本页顶部的 Data Studio 标签即可浏览,也可以直接按名字加载。
Config
行数
内容
fixtures
104
基准本体 —— 喂给模型的头部信息,以及结算后的 90 分钟赛果,七个盘口全部推导好(outcome_1x2、over_2_5、both_score、odd_total)
dossiers
2,208
简报索引 —— 46 快照 × 48… See the full description on the dataset page: https://huggingface.co/datasets/Social-AI-2026/worldcup2026.exorde-social-media-one-month-2024EgoNormia
EgoNormia: Benchmarking Physical-Social Norm Understanding
MohammadHossein Rezaei*,
Yicheng Fu*,
Phil Cuvin*,
Caleb Ziems,
Yanzhe Zhang,
Hao Zhu,
Diyi Yang,
🌎Website |
🤗 Dataset |
📄 arXiv |
📄 HF Paper
EgoNormia
EgoNormia is a challenging QA benchmark that tests VLMs' ability to reason over norms in context.
The datset consists of 1,853 physically grounded egocentric
interaction clips from Ego4D… See the full description on the dataset page: https://huggingface.co/datasets/open-social-world/EgoNormia.autolibra
AutoLibra ⚖️ Agent Metric Induction from Open-Ended Human Feedback
Hao Zhu,
Phil Cuvin,
Xinkai Yu,
Charlotte Ka Yee Yan,
Jason Zhang,
Diyi Yang,
Arxiv
AutoLibra
Dataset Organization
Annotation Format
## Contact
- Hao Zhu: zhuhao@stanford.edu
## Acknowledgement
## Citation
<!-- ```bibtex
@misc{rezaei2025egonormiabenchmarkingphysicalsocial,
title={EgoNormia: Benchmarking Physical Social Norm… See the full description on the dataset page: https://huggingface.co/datasets/open-social-world/autolibra.
