CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PKU-Alignment /BeaverTails Dataset Card for BeaverTails BeaverTails is an AI safety-focused collection comprising a series of datasets. This repository includes human-labeled data consisting of question-answer (QA) pairs, each identified with their corresponding harm categories. It should be noted that a single QA pair can be associated with more than one category. The 14 harm categories are defined as follows: Animal Abuse: This involves any form of cruelty or harm inflicted on animals, including physical… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/BeaverTails.texttext-classification100K<n<1M114 likes15k downloads3y agoHugging Face02PKU-Alignment /PKU-SafeRLHF Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. [🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset] Citation If PKU-SafeRLHF has contributed to your work, please consider citing… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF.tabulartext-generation100K<n<1M196 likes14k downloads2y agoHugging Face03nguyenvulebinh /asr-alignment Speech Recognition Alignment Dataset This dataset is a variation of several widely-used ASR datasets, encompassing Librispeech, MuST-C, TED-LIUM, VoxPopuli, Common Voice, and GigaSpeech. The difference is this dataset includes: Precise alignment between audio and text. Text that has been punctuated and made case-sensitive. Identification of named entities in the text. Usage First, install the latest version of the 🤗 Datasets package: pip install --upgrade pip pip… See the full description on the dataset page: https://huggingface.co/datasets/nguyenvulebinh/asr-alignment.audio10M<n<100M5 likes8k downloads3y agoHugging Face04nyu-visionx /Cambrian-Alignment Cambrian-Alignment Dataset Please see paper & website for more information: https://cambrian-mllm.github.io/ https://arxiv.org/abs/2406.16860 Overview Cambrian-Alignment is an question-answering alignment dataset comprised of alignment data from LLaVA, Mini-Gemini, Allava, and ShareGPT4V. Getting Started with Cambrian Alignment Data Before you start, ensure you have sufficient storage space to download and process the data. Download the Data Repository… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-Alignment.imagevisual-question-answering100K<n<1M38 likes6.9k downloads2y agoHugging Face05takuM23 /multilingual_audio_alignments Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT) A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.audioautomatic-speech-recognition10M<n<100M4 likes6.3k downloads6mo agoHugging Face06PKU-Alignment /align-anything Overview: Align-Anything Dataset A Comprehensive All-Modality Alignment Dataset with Fine-grained Preference Annotations and Language Feedback. 🏠 Homepage | 🤗 Align-Anything Dataset | 🤗 T2T_Instruction-tuning Dataset | 🤗 TI2T_Instruction-tuning Dataset | 👍 Our Official Code Repo Our world is inherently multimodal. Humans perceive the world through multiple senses, and Language Models should operate similarly. However, the development of Current Multi-Modality Foundation Models… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/align-anything.audioany-to-any10K<n<100K48 likes6.3k downloads1y agoHugging Face07AlignmentResearch /DolusChattext10K<n<100K6 likes4.5k downloads1y agoHugging Face08AlignmentResearch /soft-trigger-verifiedtext1K<n<10K0 likes4.2k downloads9mo agoHugging Face09BrainAlign /brain-lm-alignment-ds002236 Brain–language-model alignment: ds002236 (whole-brain) Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children (8.7–15.5), auditory and visual. Paper: https://pubmed.ncbi.nlm.nih.gov/31956678/ Data: https://openneuro.org/datasets/ds002236/versions/1.0.1 Generated: 2026-09-22 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number in… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds002236.documentn<1K0 likes4k downloads1h agoHugging Face10BrainAlign /brain-lm-alignment-ds006239 Brain–language-model alignment: ds006239 (whole-brain) Wang et al. 2025 — word-level phonological and semantic reading tasks in children and adolescents aged 10–17. Paper: https://www.sciencedirect.com/science/article/pii/S2352340925009692 Data: https://openneuro.org/datasets/ds006239/versions/1.0.5 Generated: 2026-09-21 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds006239.documentn<1K2 likes3.8k downloads1h agoHugging Face11Emova-ollm /emova-alignment-7m EMOVA-Alignment-7M 🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo 📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-Alignment-7M is a comprehensive dataset curated for omni-modal pre-training, including vision-language and speech-language alignment. This dataset is created using open-sourced image-text pre-training datasets, OCR datasets, and 2,000 hours of ASR and TTS data. This dataset is part of the EMOVA-Datasets… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-alignment-7m.imageimage-to-text1M<n<10M10 likes3.7k downloads2y agoHugging Face12AAdonis /multilingual_audio_alignments Multilingual MFA-Aligned Speech Dataset A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and phonemes.… See the full description on the dataset page: https://huggingface.co/datasets/AAdonis/multilingual_audio_alignments.audioautomatic-speech-recognition10M<n<100M27 likes3.5k downloads5mo agoHugging Face13cfierro /alignment_faking_claude_completionstext1K<n<10K0 likes2.8k downloads1y agoHugging Face14FabienRoger /alignment_faking_harm_answerstext1K<n<10K0 likes2.8k downloads1y agoHugging Face15gilkeyio /librispeech-alignments Dataset Card for Librispeech Alignments Librispeech with alignments generated by the Montreal Forced Aligner. The original alignments in TextGrid format can be found here Dataset Details Dataset Description Librispeech is a corpus of read English speech, designed for training and evaluating automatic speech recognition (ASR) systems. The dataset contains 1000 hours of 16kHz read English speech derived from audiobooks. The Montreal Forced Aligner (MFA) was used… See the full description on the dataset page: https://huggingface.co/datasets/gilkeyio/librispeech-alignments.audioautomatic-speech-recognition100K<n<1M21 likes2.4k downloads3y agoHugging Face16Project-AgML /Agri-LLaVA_Agricultural_Pests_And_Diseases_Feature_Alignment_Dataset Agri-LLaVA Agri-LLaVA is a large multimodal instruction dataset for agriculture, pairing crop/leaf images with multi-turn diagnostic conversations about plant diseases, pests, and nutrient deficiencies. It is compiled from 16 public source datasets (see the license table below). This dataset has been converted to Parquet format with image bytes embedded directly, standardized to the HF image_text_to_text format with a single conversational messages schema. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/Agri-LLaVA_Agricultural_Pests_And_Diseases_Feature_Alignment_Dataset.imageimage-text-to-text100K<n<1M0 likes2.3k downloads2mo agoHugging Face17BrainAlign /brain-lm-alignment-ds001894 Brain–language-model alignment: ds001894 (whole-brain) Lytle et al. 2019 — longitudinal word-level phonological processing in children scanned twice, at roughly 10 and 12 years old. Paper: https://www.nature.com/articles/s41597-019-0338-5 Data: https://openneuro.org/datasets/ds001894/versions/1.4.2 Generated: 2026-09-21 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds001894.documentn<1K0 likes2k downloads7h agoHugging Face18AlignmentResearch /ClearHarmtext1K<n<10K4 likes1.9k downloads1y agoHugging Face19PKU-Alignment /MM-SafetyBenchWarning: This dataset may contain sensitive or harmful content. Users are advised to handle it with care and ensure that their use complies with relevant ethical guidelines and legal requirements. Usage and License Notices: The dataset is intended and licensed for research use only. They are also restricted to uses that follow the license agreement GPT-4 and Stable Diffusion. The dataset is CC BY NC 4.0 (allowing only non-commercial use). Data Source: For more information about the dataset… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/MM-SafetyBench.image1K<n<10K8 likes1.9k downloads2y agoHugging Face20PKU-Alignment /PKU-SafeRLHF-10K Paper You can find more information in our paper. Dataset Paper: https://arxiv.org/abs/2307.04657 tabulartext-generation10K<n<100K62 likes1.6k downloads3y agoHugging Face21HannahRoseKirk /prism-alignment Dataset Card for PRISM PRISM is a diverse human feedback dataset for preference and value alignment in Large Language Models (LLMs). It maps the characteristics and stated preferences of humans from a detailed survey onto their real-time interactions with LLMs and contextual preference ratings Dataset Details There are two sequential stages: first, participants complete a Survey where they answer questions about their demographics and stated preferences, then proceed to… See the full description on the dataset page: https://huggingface.co/datasets/HannahRoseKirk/prism-alignment.tabular10K<n<100K105 likes1.5k downloads2y agoHugging Face22Anthropic /alignment-faking-rl Transcripts from Towards training-time mitigations for alignment faking in RL This dataset contains the full evaluation transcripts through the RL runs for all model organisms in our blog post, Towards training-time mitigations for alignment faking in RL. Each file in encrypted_transcripts/ corresponds to one RL training run. Precautions against pretraining data poisoning In order to avoid our model organisms' misaligned reasoning from accidentally appearing in… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/alignment-faking-rl.tabular1M<n<10M18 likes1.3k downloads9mo agoHugging Face23PKU-Alignment /PKU-SafeRLHF-30K Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. Dataset Summary The preference dataset consists of 30k+ expert comparison data. Each entry in this dataset includes two responses to a question, along with safety… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-30K.tabulartext-generation10K<n<100K14 likes1.2k downloads3y agoHugging Face24nace-ai /policy-alignment-verification-dataset Policy Alignment Verification Dataset 🌐 NAVI's Ecosystem 🌐 🌍 NAVI Platform – Dive into NAVI's full capabilities and explore how it ensures policy alignment and compliance. 🤗 NAVI-small-preview – Access the open-weights version of NAVI designed for policy verification. 📜 API Docs – Your starting point for integrating NAVI into your applications. 📝 Blogpost: Policy-Driven Safeguards Comparison – A deep dive into the challenges and solutions NAVI addresses. ✨… See the full description on the dataset page: https://huggingface.co/datasets/nace-ai/policy-alignment-verification-dataset.texttext-classificationn<1K4 likes845 downloads2y agoHugging Face25jcnf /targeting-alignment Dataset Card The datasets in this repository correspond to the embeddings used in "Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs". For each model, source dataset (input prompts) and setting (benign or adversarial), the corresponding dataset contains the base input prompt, the (deterministic) output of the model, the representations of the input at each layer of the model and the corresponding unsafe/safe labels (1 for unsafe, 0 for safe). Dataset… See the full description on the dataset page: https://huggingface.co/datasets/jcnf/targeting-alignment.tabulartext-generation1M<n<10M0 likes845 downloads2y agoHugging Face26AIM-Intelligence /COMPASS-Policy-Alignment-Testbed-Dataset COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs This dataset evaluates how well Large Language Models (LLMs) follow organization-specific policies in realistic enterprise-style settings. What is COMPASS? COMPASS is a framework for evaluating policy alignment: given only an organization’s policy (e.g., allow/deny rules), it enables you to benchmark whether an LLM’s responses comply with that policy in structured, enterprise-like… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/COMPASS-Policy-Alignment-Testbed-Dataset.texttext-generation1K<n<10K12 likes746 downloads24d agoHugging Face27Rapidata /Flux_SD3_MJ_Dalle_Human_Alignment_Dataset NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Alignment_Dataset Rapidata Image Generation Alignment Dataset This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment. Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset Link to the Preference dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Alignment_Dataset.imagetext-to-image10K<n<100K16 likes643 downloads2y agoHugging Face28PKU-Alignment /BeaverTails-Evaluation Dataset Card for BeaverTails-Evaluation BeaverTails is an AI safety-focused collection comprising a series of datasets. This repository contains test prompts specifically designed for evaluating language model safety. It is important to note that although each prompt can be connected to multiple categories, only one category is labeled for each prompt. The 14 harm categories are defined as follows: Animal Abuse: This involves any form of cruelty or harm inflicted on animals… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/BeaverTails-Evaluation.texttext-classificationn<1K15 likes632 downloads3y agoHugging Face29facebook /community-alignment-dataset Community Alignment Github   |   Paper Dataset Community Alignment is a large-scale open source, multilingual and multi-turn preference dataset to align LLMs with human preferences across cultures. Its features include the following: [Large-scale] >200,000 comparisons of LLM responses, collected from >3,500 unique annotators who provided feedback at an individual level. [Multilingual] Contains comparisons in English, French, Italian, Hindi, and Portuguese. 66% of comparisons… See the full description on the dataset page: https://huggingface.co/datasets/facebook/community-alignment-dataset.tabular10K<n<100K42 likes597 downloads7mo agoHugging Face30Alignment-Lab-AI /Open-Web-Math Keiran Paster*, Marco Dos Santos*, Zhangir Azerbayev, Jimmy Ba GitHub | ArXiv | PDF OpenWebMath is a dataset containing the majority of the high-quality, mathematical text from the internet. It is filtered and extracted from over 200B HTML files on Common Crawl down to a set of 6.3 million documents containing a total of 14.7B tokens. OpenWebMath is intended for use in pretraining and finetuning large language models. You can download the dataset using Hugging Face: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/Open-Web-Math.text1M<n<10M5 likes583 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.