CoolFace
20 results

social

allenai /social_i_qaWe introduce Social IQa: Social Interaction QA, a new question-answering benchmark for testing social commonsense intelligence. Contrary to many prior benchmarks that focus on physical or taxonomic knowledge, Social IQa focuses on reasoning about people’s actions and their social implications. For example, given an action like "Jesse saw a concert" and a question like "Why did Jesse do this?", humans can easily infer that Jesse wanted "to see their favorite performer" or "to enjoy the music", and not "to see what's happening inside" or "to see if it works". The actions in Social IQa span a wide variety of social situations, and answer candidates contain both human-curated answers and adversarially-filtered machine-generated candidates. Social IQa contains over 37,000 QA pairs for evaluating models’ abilities to reason about the social implications of everyday events and situations. (Less)31 likes64k downloads10mo agoHugging Faceruggsea /social-sim-bench-genstext1K<n<10K0 likes7.9k downloads3mo agoHugging FaceSocial-AI-2026 /worldcup2026 ⚽ WorldCup Arena A Leakage-Free Forecasting Benchmark on a Live Tournament Can a language model forecast a match — when the match had not been played at the moment it was asked? &nbsp; 🌐 &nbsp; Language / 语言 &nbsp;:&nbsp; 中文 &nbsp; ▾ &nbsp; 📊 四张表 点开本页顶部的 Data Studio 标签即可浏览,也可以直接按名字加载。 Config 行数 内容 fixtures 104 基准本体 —— 喂给模型的头部信息,以及结算后的 90 分钟赛果,七个盘口全部推导好(outcome_1x2、over_2_5、both_score、odd_total) dossiers 2,208 简报索引 —— 46 快照 × 48… See the full description on the dataset page: https://huggingface.co/datasets/Social-AI-2026/worldcup2026.tabularquestion-answering1K<n<10K0 likes6.2k downloads2mo agoHugging FaceExorde /exorde-social-media-one-month-2024texttext-classification100M<n<1B29 likes2.9k downloads2y agoHugging Faceopen-social-world /EgoNormia EgoNormia: Benchmarking Physical-Social Norm Understanding MohammadHossein Rezaei*,  Yicheng Fu*,  Phil Cuvin*,  Caleb Ziems,  Yanzhe Zhang,  Hao Zhu,  Diyi Yang,  🌎Website | 🤗 Dataset | 📄 arXiv | 📄 HF Paper EgoNormia EgoNormia is a challenging QA benchmark that tests VLMs' ability to reason over norms in context. The datset consists of 1,853 physically grounded egocentric interaction clips from Ego4D… See the full description on the dataset page: https://huggingface.co/datasets/open-social-world/EgoNormia.imagevisual-question-answering1K<n<10K7 likes2.8k downloads1y agoHugging Faceopen-social-world /autolibra AutoLibra ⚖️ Agent Metric Induction from Open-Ended Human Feedback Hao Zhu,  Phil Cuvin,  Xinkai Yu,  Charlotte Ka Yee Yan,  Jason Zhang,  Diyi Yang,  Arxiv AutoLibra Dataset Organization Annotation Format ## Contact - Hao Zhu: zhuhao@stanford.edu ## Acknowledgement ## Citation <!-- ```bibtex @misc{rezaei2025egonormiabenchmarkingphysicalsocial, title={EgoNormia: Benchmarking Physical Social Norm… See the full description on the dataset page: https://huggingface.co/datasets/open-social-world/autolibra.1 likes1.7k downloads6mo agoHugging Face