datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multi-IF
Dataset Summary
We introduce Multi-IF, a new benchmark designed to assess LLMs' proficiency in following multi-turn and multilingual instructions. Multi-IF, which utilizes a hybrid framework combining LLM and human annotators, expands upon the IFEval by incorporating multi-turn sequences and translating the English prompts into another 7 languages, resulting in a dataset of 4501 multilingual conversations, where each has three turns. Our evaluation of 14 state-of-the-art LLMs on… See the full description on the dataset page: https://huggingface.co/datasets/facebook/Multi-IF.ExploreToM
Data sample for ExploreToM: Program-guided adversarial data generation for theory of mind reasoning
ExploreToM is the first framework to allow large-scale generation of diverse and challenging theory of mind data for robust training and evaluation.
Our approach leverages an A* search over a custom domain-specific language to produce complex story structures and novel, diverse, yet plausible scenarios to stress test the limits of LLMs.
Our A* search procedure aims to find… See the full description on the dataset page: https://huggingface.co/datasets/facebook/ExploreToM.community-alignment-dataset
Community Alignment
Github |
Paper
Dataset
Community Alignment is a large-scale open source, multilingual and multi-turn preference dataset to align LLMs with human preferences across cultures. Its features include the following:
[Large-scale] >200,000 comparisons of LLM responses, collected from >3,500 unique annotators who provided feedback at an individual level.
[Multilingual] Contains comparisons in English, French, Italian, Hindi, and Portuguese. 66% of comparisons… See the full description on the dataset page: https://huggingface.co/datasets/facebook/community-alignment-dataset.EgoAVU_data
[CVPR2026 HIGHLIGHT] EgoAVU, [ICASSP2026 Oral] Exploring Audio Hallucination in Egocentric Video Understanding
Official Implementation of EgoAVU: Egocentric Audio-Visual Understanding and Exploring Audio Hallucination in Egocentric Video Understanding
See our github for the code and setup instructions.
Check out our homepage, paper (CVPR) and paper (ICASSP) for more information.
We introduce EgoAVU, a scalable and automated data engine to enable egocentric audio–visual… See the full description on the dataset page: https://huggingface.co/datasets/facebook/EgoAVU_data.OMC25
Open Molecular Crystals 2025 (OMC25)
Dataset
Dataset
LICENSE: The OMC25 dataset is provided under a CC-BY-4.0 license
OMC25 represents the largest high quality molecular crystal DFT dataset.
OMC25 was generated at the PBE-D3 level of theory as implemented in Vienna Ab initio Simulation Package (VASP). OMC25 includes structures sampled from relaxation trajectories of molecular crystals generated by Genarris 3.0 starting from molecules in the OE62 dataset.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/OMC25.facebook-personality-recognition-wcpr13The Workshop on Computational Personality Recognition 2013 was a competition based on this Facebook dataset.
The purpose is to predict the personality scores or classes from text and ego-network data
reference paper: https://ojs.aaai.org/index.php/ICWSM/article/view/14467/14316
Y-NQ
Dataset Card for Y-NQ
The dataset is available in this csv file. The dataset is licensed under the Apache 2.0 license
Dataset Description
Question ID: Unique identifier from Natural Question
Split: Training or validation split from Natural Question
English Document: English text document
English Question: Question in English
English Long Answer: Detailed answer in English
English Short Answer: Brief answer in English
Yorùbá Document: Yorùbá text document
Yorùbá… See the full description on the dataset page: https://huggingface.co/datasets/facebook/Y-NQ.facebook_spam_detection
Facebook Spam Detection Dataset
Dataset Summary
This dataset contains 600 Facebook profiles with behavioral and activity features designed for spam detection in social media. The dataset enables binary classification to distinguish between spam accounts (Label=1) and legitimate accounts (Label=0), providing insights into spammer behavior patterns on Facebook.
Dataset Details
Total Samples: 600 profiles
Classes: Binary (0 = Legitimate, 1 = Spam)
Class… See the full description on the dataset page: https://huggingface.co/datasets/nahiar/facebook_spam_detection.beyond_the_lab_neurips_paperSCRuB-dataset
SCRuB — Social Concept Reasoning under Rubric-Based Evaluation
SCRuB is a dataset suite for studying how large language models handle socially sensitive, open-ended essay prompts. It comprises three components:
Component
Description
Rows
SCRuBSample
30 curated study prompts used as stimuli in a human annotation study
30
SCRuBAnnotations
Expert essays, model responses, and quality judgments from a two-task annotation study
300 + 78 + 20 + 900 + 900
SCRuBEval4,711… See the full description on the dataset page: https://huggingface.co/datasets/facebook/SCRuB-dataset.FacebookDecadeCorporaFacebook-datasetfacebook-users-dataFacebook Users Engagement Analysis
Author: Tamara Banaim
Dataset: Pseudo Facebook Dataset (Kaggle, uploaded to Hugging Face)
Overview-
This project analyzes data from 99,003 Facebook users, focusing on demographic information and engagement metrics such as likes given, likes received, friend count, and account tenure. The analysis explores how age and user activity are related, and what factors influence engagement on the platform.
Objective-
To examine how age and… See the full description on the dataset page: https://huggingface.co/datasets/tamarabanaim/facebook-users-data.
