CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nyu-visionx /Cambrian-Alignment Cambrian-Alignment Dataset Please see paper & website for more information: https://cambrian-mllm.github.io/ https://arxiv.org/abs/2406.16860 Overview Cambrian-Alignment is an question-answering alignment dataset comprised of alignment data from LLaVA, Mini-Gemini, Allava, and ShareGPT4V. Getting Started with Cambrian Alignment Data Before you start, ensure you have sufficient storage space to download and process the data. Download the Data Repository… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-Alignment.imagevisual-question-answering100K<n<1M38 likes6.5k downloads2y agoHugging Face02Rapidata /human-alignment-preferences-images Rapidata Image Generation Alignment Dataset This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please consider liking it. Overview One of the largest human annotated alignment datasets for text-to-image models, this release contains over 1,200,000 human… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-alignment-preferences-images.imagetext-to-image10K<n<100K17 likes704 downloads2y agoHugging Face03PKU-Alignment /DeceptionBench DeceptionBench: A Comprehensive Benchmark for Evaluating Deceptive Behaviors in Large Language Models 🔍 Overview DeceptionBench is the first systematic benchmark designed to assess deceptive behaviors in Large Language Models (LLMs). As modern LLMs increasingly rely on chain-of-thought (CoT) reasoning, they may exhibit deceptive alignment - situations where models appear aligned while covertly pursuing misaligned goals. This benchmark addresses a critical gap in AI… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/DeceptionBench.texttext-classificationn<1K4 likes338 downloads1y agoHugging Face04daiandy /task-alignment-datasetRelease version: (2026-07-16) Three benchmarks for evaluating LLM task alignment under underspecification. Each row is one task specification: the assistant must interact to identify the user's ground-truth task x* from a fixed set of 15 candidate specifications, given only an evolving natural-language intent from a user simulator. Files Dataset File Rows GDPVal (knowledge-work tasks) gdpval_v4_camera_ready_n88.csv 88 Terminal-Bench (coding tasks)… See the full description on the dataset page: https://huggingface.co/datasets/daiandy/task-alignment-dataset.texttext-classificationn<1K0 likes167 downloads2mo agoHugging Face05URSA-MATH /URSA_Alignment_860K URSA_Alignment_860K This dataset is used for the vision-language alignment phase of training the URSA-7B model. Image data can be downloaded from the following address: MAVIS: https://github.com/ZrrSkywalker/MAVIS, https://drive.google.com/drive/folders/1LGd2JCVHi1Y6IQ7l-5erZ4QRGC4L7Nol. Multimath: https://huggingface.co/datasets/pengshuai-rin/multimath-300k. Geo170k: https://huggingface.co/datasets/Luckyjhg/Geo170K. The image data in the MMathCoT-1M dataset is still available.… See the full description on the dataset page: https://huggingface.co/datasets/URSA-MATH/URSA_Alignment_860K.textquestion-answering100K<n<1M7 likes81 downloads2y agoHugging Face06projecte-aina /hhh_alignment_ca Dataset Card for hhh_alignment_ca hhh_alignment_ca is a question answering dataset in Catalan, professionally translated from the main version of the hhh_alignment dataset in English. Dataset Details Dataset Description hhh_alignment_ca (Helpful, Honest, & Harmless - a Pragmatic Alignment Evaluation - Catalan) is designed to evaluate language models on alignment, pragmatically broken down into the categories of helpfulness, honesty/accuracy, harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/hhh_alignment_ca.textquestion-answeringn<1K0 likes52 downloads2y agoHugging Face07BSC-LT /hhh_alignment_es Dataset Card for hhh_alignment_es hhh_alignment_es is a question answering dataset in Spanish, professionally translated from the main version of the hhh_alignment dataset in English. Dataset Details Dataset Description hhh_alignment_es (Helpful, Honest, & Harmless - a Pragmatic Alignment Evaluation - Spanish) is designed to evaluate language models on alignment, pragmatically broken down into the categories of helpfulness, honesty/accuracy, harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/hhh_alignment_es.texttext-classificationn<1K0 likes34 downloads2y agoHugging Face08PKU-Alignment /self-monitor Self-Monitor Dataset This dataset contains supervised fine-tuning (SFT) data used in the research paper "Mitigating Deceptive Alignment via Self-Monitoring" (arXiv:2505.18807). Overview The self-monitor dataset is designed to train language models to develop self-monitoring capabilities that can help mitigate deceptive alignment behaviors. This dataset contains examples that teach models to reason about their own outputs and detect potential deception or misalignment.… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/self-monitor.tabulartext-generation10K<n<100K0 likes28 downloads1y agoHugging Face09IMoonKeyBoy /PKU-Alignment-Graphtabularquestion-answering10K<n<100K0 likes28 downloads9mo agoHugging Face10robworks-software /ccisd-teks-alignment-split [!WARNING] Deprecated - use ccisd-teks-alignment instead. This dataset is superseded: the two contain the same 428 rows with the same 12 columns; this copy only adds a train/validation/test partition, which you can reproduce in one line. Nothing here is unique to it. It stays online so existing references keep resolving, but it will not be updated. New work should point at robworks-software/ccisd-teks-alignment. CCISD TEKS Alignment (pre-split) The same 428 TEKS-to-course… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/ccisd-teks-alignment-split.texttext-classificationn<1K0 likes24 downloads2mo agoHugging Face11ustc-zhangzm /trustworthy-alignment Trustworthy Alignment of Retrieval-Augmented Large Language Models via Reinforcement Learning Official repository for Trustworthy Alignment of Retrieval-Augmented Large Language Models via Reinforcement Learning GitHub Repository: https://github.com/zmzhang2000/trustworthy-alignment HuggingFace Hub: https://huggingface.co/datasets/ustc-zhangzm/trustworthy-alignment Paper: https://proceedings.mlr.press/v235/zhang24bg.html Usage from datasets importload_dataset… See the full description on the dataset page: https://huggingface.co/datasets/ustc-zhangzm/trustworthy-alignment.textquestion-answering10K<n<100K1 likes21 downloads2y agoHugging Face12EpistemeAI /EpistemeAI-alignment-safety-40-chat Dataset Card for EpistemeAI-alignement-safety-40-chat This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/EpistemeAI/EpistemeAI-alignement-safety-40-chat/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/EpistemeAI/EpistemeAI-alignment-safety-40-chat.texttext-generationn<1K0 likes16 downloads1y agoHugging Face13August4293 /Self_Alignment_Preference-Dataset Mistral Self-Alignment Preference Dataset Warning: This dataset contains harmful and offensive data! Proceed with caution. The Mistral Self-Alignment Preference Dataset was generated by Mistral 7b using the Anthropics Red Teaming Prompts dataset available at Hugging Face - Anthropics Red Teaming Prompts Dataset. The data generation process utilized the Preference Data Generation Notebook, which can be found here. The purpose of this dataset is to facilitate self-alignment, as… See the full description on the dataset page: https://huggingface.co/datasets/August4293/Self_Alignment_Preference-Dataset.texttext-generation1K<n<10K0 likes13 downloads3y agoHugging Face14gsoisson /alignment-internship-exercise Dataset Card for the Alignement Internship Exercise Dataset Description This dataset provides a list of questions accompanied by Phi-2's best answer to them, as ranked by OpenAssitant's reward model. Dataset Creation The questions were handpicked from the LDJnr/Capybara, Open-Orca/OpenOrca and truthful_qa datasets, the coding exercise is from LeetCode's top 100 liked questions and I found the last prompt on a blog and modified it. I have chosen these prompts… See the full description on the dataset page: https://huggingface.co/datasets/gsoisson/alignment-internship-exercise.textquestion-answeringn<1K0 likes10 downloads3y agoHugging Face15MatinaAI /alignment_datasetsgated 🧠 Persian Cultural Alignment Dataset for LLMs This repository contains a high-quality, Alignment dataset for cultural alignment of large language models (LLMs) in the Persian language. The dataset is curated using hybrid strategies that incorporate culturally grounded generation, multi-turn dialogues, translation, and augmentation methods, making it suitable for SFT, DPO, RLHF, and alignment evaluation. 📚 Dataset Overview Domain Methods Used Culinary… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/alignment_datasets.textquestion-answering10K<n<100K0 likes10 downloads1y agoHugging Face16FlameF0X /Safety_Alignment_Benchmarkgatedtexttext-generationn<1K0 likes9 downloads10mo agoHugging Face17robworks-software /ccisd-teks-alignment CCISD TEKS Alignment 428 Texas Essential Knowledge and Skills (TEKS) student expectations mapped to 25 Clear Creek ISD high school courses, with STAAR-tested status flagged. Loading from datasets import load_dataset ds = load_dataset("robworks-software/ccisd-teks-alignment") 428 rows, single train split. A pre-split version of the same 428 rows is published as ccisd-teks-alignment-split. Contents 428 distinct TEKS codes (e.g. ELAR.9.1.A), each… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/ccisd-teks-alignment.texttext-classificationn<1K0 likes8 downloads2mo agoHugging Face18MichiganNLP /misfired-alignmentgated VETO: A Benchmark for Misfired Alignment ⚠️ Content warning. This dataset references historically stereotyped demographic groups and contains potentially disturbing content, included only to measure a failure mode in LLMs. It is not an endorsement of any stereotype, and these findings are not an argument against alignment. VETO accompanies the paper "The Wrong Kind of Right: Quantifying and Localizing Misfired Alignment in LLMs." It measures misfired alignment — when an… See the full description on the dataset page: https://huggingface.co/datasets/MichiganNLP/misfired-alignment.textquestion-answering1K<n<10K0 likes7 downloads3mo agoHugging Face19Skywalker-Harrison-mbz /cultural_alignment_ar_en_dpotextquestion-answering10K<n<100K0 likes4 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.