CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01facebook /natural_reasoningNaturalReasoning is a large-scale dataset for general reasoning tasks. It consists of high-quality challenging reasoning questions backtranslated from pretraining corpora DCLM and FineMath. The questions have been deduplicated and decontaminated from popular reasoning benchmarks including MATH, GPQA, MMLU-Pro, MMLU-STEM. For each question, we extract the reference final answer from the original document from the pretraining corpora if possible. We also provide a model-generated response from… See the full description on the dataset page: https://huggingface.co/datasets/facebook/natural_reasoning.texttext-generation1M<n<10M585 likes2.7k downloads2y agoHugging Face02facebook /cyberseceval3-visual-prompt-injection Dataset Card for CyberSecEval 3 - Visual Prompt Injection Benchmark Dataset Details Dataset Description This dataset provides a multimodal benchmark for visual prompt injection, with text/image inputs. It is part of CyberSecEval 3, the third edition of Meta's flagship suite of security benchmarks for LLMs to measure cybersecurity risks and capabilities across multiple domains. Language(s): English License: MIT Dataset Sources Repository: Link… See the full description on the dataset page: https://huggingface.co/datasets/facebook/cyberseceval3-visual-prompt-injection.imagetext-generation1K<n<10K10 likes2.6k downloads2y agoHugging Face03facebook /BigOBench 👋 Overview 🚀 Introduction 📋 Getting Started with the data 🔥 problem_and_human_solutions_list.jsonl 🔥 complexity_labels_light.jsonl 🔥 complexity_labels_full.jsonl 🔥 time_complexity_test_set.jsonl 🔥 space_complexity_test_set.jsonl License 📝 Citation 🚀 Introduction BigO(Bench) is a benchmark of ~300 code problems to be solved in Python, along with 3,105 coding problems… See the full description on the dataset page: https://huggingface.co/datasets/facebook/BigOBench.tabulartext-classification1M<n<10M9 likes851 downloads2y agoHugging Face04facebook /multilokogated MultiLoKo: a multilingual local knowledge benchmark for LLMs MultiLoKo is a multilingual knowledge benchmark, covering 30 languages plus English. The questions are separately sourced for each language, with an annotation protocol designed to target locally relevant topics for the respective language. MultiLoKo contains the original data for each language, as well as both human and machine-authored translations of each non-English subset into English and vice versa, facilitating… See the full description on the dataset page: https://huggingface.co/datasets/facebook/multiloko.textquestion-answering10K<n<100K7 likes849 downloads1y agoHugging Face05facebook /llamafirewall-alignmentcheck-evals Dataset Card for LlamaFirewall AlignmentCheck Evals Dataset Details Dataset Description This dataset provides a dataset for prompt injection in an agentic environment. It is part of LlamaFirewall, an open-source security focused guardrail framework designed to serve as a final layer of defense against security risks associated with AI Agents. Specifically, this dataset is designed to evaluate the susceptibility of language models, and detect any misalignment… See the full description on the dataset page: https://huggingface.co/datasets/facebook/llamafirewall-alignmentcheck-evals.tabulartext-generation1K<n<10K4 likes123 downloads1y agoHugging Face06abdelhaqueidali /Amazigh-Facebook-Data-Export Dataset Card for Amazigh Facebook Dataset (Filtered) This dataset contains a curated, filtered collection of personal Facebook posts and comments written by an Amazigh speaker. It captures real-world communication usage of Amazigh in the Amazigh script (Tifinagh). The speaker utilized Southern Moroccan Amazigh (Tashelhit) and leaned to the extent possible toward Standard Moroccan Amazigh. Some posts and comments uniquely contain parallel multi-script and multi-language segments… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Amazigh-Facebook-Data-Export.texttext-generation10K<n<100K0 likes28 downloads4mo agoHugging Face07facells /seshat-perspectiveSeshat-perspective is the first historical databank synthetically annotated with a perspectivist approach by means of multiple Large Language Models: Deepseek (dr1), Llama (l31l) and Mistral (m3m) Paper with field description and validation procedure: https://github.com/facells/fabio-celli-publications/blob/main/docs/2026_perspective_seshat_clicit26.pdf Code for replication: https://colab.research.google.com/drive/1_4aUNGjl7_uhLZZKE7mAYHPhWYUZ9jvr?usp=sharing tabulartext-generation1K<n<10K0 likes24 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.