datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flores
Dataset Card for Flores 200
Dataset Summary
⚠️ This repository is no longer being updated ⚠️
A newer version of the FLORES dataset managed by the Open Language Data Initiative
is available at https://huggingface.co/datasets/openlanguagedata/flores_plus.
FLORES is a benchmark dataset for machine translation between English and low-resource languages.
The creation of FLORES-200 doubles the existing language coverage of FLORES-101.
Given the nature of the new… See the full description on the dataset page: https://huggingface.co/datasets/facebook/flores.BigOBench
👋 Overview
🚀 Introduction
📋 Getting Started with the data
🔥 problem_and_human_solutions_list.jsonl
🔥 complexity_labels_light.jsonl
🔥 complexity_labels_full.jsonl
🔥 time_complexity_test_set.jsonl
🔥 space_complexity_test_set.jsonl
License
📝 Citation
🚀 Introduction
BigO(Bench) is a benchmark of ~300 code problems to be solved in Python, along with 3,105 coding problems… See the full description on the dataset page: https://huggingface.co/datasets/facebook/BigOBench.llamafirewall-alignmentcheck-evals
Dataset Card for LlamaFirewall AlignmentCheck Evals
Dataset Details
Dataset Description
This dataset provides a dataset for prompt injection in an agentic environment. It is part of LlamaFirewall, an open-source security focused guardrail framework designed to serve as a final layer of defense against security risks associated with AI Agents. Specifically, this dataset is designed to evaluate the susceptibility of language models, and detect any misalignment… See the full description on the dataset page: https://huggingface.co/datasets/facebook/llamafirewall-alignmentcheck-evals.seshat-perspectiveSeshat-perspective is the first historical databank synthetically annotated with a perspectivist approach by means of multiple Large Language Models: Deepseek (dr1), Llama (l31l) and Mistral (m3m)
Paper with field description and validation procedure: https://github.com/facells/fabio-celli-publications/blob/main/docs/2026_perspective_seshat_clicit26.pdf
Code for replication: https://colab.research.google.com/drive/1_4aUNGjl7_uhLZZKE7mAYHPhWYUZ9jvr?usp=sharing
SCRuB-dataset
SCRuB — Social Concept Reasoning under Rubric-Based Evaluation
SCRuB is a dataset suite for studying how large language models handle socially sensitive, open-ended essay prompts. It comprises three components:
Component
Description
Rows
SCRuBSample
30 curated study prompts used as stimuli in a human annotation study
30
SCRuBAnnotations
Expert essays, model responses, and quality judgments from a two-task annotation study
300 + 78 + 20 + 900 + 900
SCRuBEval4,711… See the full description on the dataset page: https://huggingface.co/datasets/facebook/SCRuB-dataset.
