datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
airs-bench
AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents
The AI Research Science Benchmark (AIRS-Bench) quantifies the autonomous research abilities of LLM agents in the area of machine learning. AIRS-Bench comprises 20 tasks from state-of-the-art machine learning papers spanning diverse domains: NLP, Code, Math, biochemical modelling, and time series forecasting.
Each task is specified by a ⟨problem, dataset, metric⟩ triplet and a SOTA value. The agent receives the… See the full description on the dataset page: https://huggingface.co/datasets/facebook/airs-bench.HoneyBee
HoneyBee: Data Recipes for Vision-Language Reasoners
This is the official data release for the paper: https://arxiv.org/abs/2510.12225.
Github Repo: https://github.com/facebookresearch/HoneyBee_VLM.
Abstract
Recent advances in vision-language models (VLMs) have made them highly effective at reasoning tasks. However, the principles underlying the construction of performant VL reasoning training datasets remain poorly understood. In this work, we introduce several data… See the full description on the dataset page: https://huggingface.co/datasets/facebook/HoneyBee.SCRuB-dataset
SCRuB — Social Concept Reasoning under Rubric-Based Evaluation
SCRuB is a dataset suite for studying how large language models handle socially sensitive, open-ended essay prompts. It comprises three components:
Component
Description
Rows
SCRuBSample
30 curated study prompts used as stimuli in a human annotation study
30
SCRuBAnnotations
Expert essays, model responses, and quality judgments from a two-task annotation study
300 + 78 + 20 + 900 + 900
SCRuBEval4,711… See the full description on the dataset page: https://huggingface.co/datasets/facebook/SCRuB-dataset.
