CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PKU-Alignment /BeaverTails Dataset Card for BeaverTails BeaverTails is an AI safety-focused collection comprising a series of datasets. This repository includes human-labeled data consisting of question-answer (QA) pairs, each identified with their corresponding harm categories. It should be noted that a single QA pair can be associated with more than one category. The 14 harm categories are defined as follows: Animal Abuse: This involves any form of cruelty or harm inflicted on animals, including physical… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/BeaverTails.texttext-classification100K<n<1M114 likes15k downloads3y agoHugging Face02PKU-Alignment /PKU-SafeRLHF Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. [🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset] Citation If PKU-SafeRLHF has contributed to your work, please consider citing… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF.tabulartext-generation100K<n<1M196 likes14k downloads2y agoHugging Face03nguyenvulebinh /asr-alignment Speech Recognition Alignment Dataset This dataset is a variation of several widely-used ASR datasets, encompassing Librispeech, MuST-C, TED-LIUM, VoxPopuli, Common Voice, and GigaSpeech. The difference is this dataset includes: Precise alignment between audio and text. Text that has been punctuated and made case-sensitive. Identification of named entities in the text. Usage First, install the latest version of the 🤗 Datasets package: pip install --upgrade pip pip… See the full description on the dataset page: https://huggingface.co/datasets/nguyenvulebinh/asr-alignment.audio10M<n<100M5 likes8k downloads3y agoHugging Face04nyu-visionx /Cambrian-Alignment Cambrian-Alignment Dataset Please see paper & website for more information: https://cambrian-mllm.github.io/ https://arxiv.org/abs/2406.16860 Overview Cambrian-Alignment is an question-answering alignment dataset comprised of alignment data from LLaVA, Mini-Gemini, Allava, and ShareGPT4V. Getting Started with Cambrian Alignment Data Before you start, ensure you have sufficient storage space to download and process the data. Download the Data Repository… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-Alignment.imagevisual-question-answering100K<n<1M38 likes6.9k downloads2y agoHugging Face05takuM23 /multilingual_audio_alignments Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT) A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.audioautomatic-speech-recognition10M<n<100M4 likes6.3k downloads6mo agoHugging Face06PKU-Alignment /align-anything Overview: Align-Anything Dataset A Comprehensive All-Modality Alignment Dataset with Fine-grained Preference Annotations and Language Feedback. 🏠 Homepage | 🤗 Align-Anything Dataset | 🤗 T2T_Instruction-tuning Dataset | 🤗 TI2T_Instruction-tuning Dataset | 👍 Our Official Code Repo Our world is inherently multimodal. Humans perceive the world through multiple senses, and Language Models should operate similarly. However, the development of Current Multi-Modality Foundation Models… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/align-anything.audioany-to-any10K<n<100K48 likes6.3k downloads1y agoHugging Face07AlignmentResearch /DolusChattext10K<n<100K6 likes4.5k downloads1y agoHugging Face08AlignmentResearch /soft-trigger-verifiedtext1K<n<10K0 likes4.2k downloads9mo agoHugging Face09BrainAlign /brain-lm-alignment-ds002236 Brain–language-model alignment: ds002236 (whole-brain) Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children (8.7–15.5), auditory and visual. Paper: https://pubmed.ncbi.nlm.nih.gov/31956678/ Data: https://openneuro.org/datasets/ds002236/versions/1.0.1 Generated: 2026-09-21 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number in… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds002236.documentn<1K0 likes4k downloads1h agoHugging Face10BrainAlign /brain-lm-alignment-ds006239 Brain–language-model alignment: ds006239 (whole-brain) Wang et al. 2025 — word-level phonological and semantic reading tasks in children and adolescents aged 10–17. Paper: https://www.sciencedirect.com/science/article/pii/S2352340925009692 Data: https://openneuro.org/datasets/ds006239/versions/1.0.5 Generated: 2026-09-21 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds006239.documentn<1K2 likes3.8k downloads2h agoHugging Face11Emova-ollm /emova-alignment-7m EMOVA-Alignment-7M 🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo 📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-Alignment-7M is a comprehensive dataset curated for omni-modal pre-training, including vision-language and speech-language alignment. This dataset is created using open-sourced image-text pre-training datasets, OCR datasets, and 2,000 hours of ASR and TTS data. This dataset is part of the EMOVA-Datasets… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-alignment-7m.imageimage-to-text1M<n<10M10 likes3.7k downloads2y agoHugging Face12AAdonis /multilingual_audio_alignments Multilingual MFA-Aligned Speech Dataset A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and phonemes.… See the full description on the dataset page: https://huggingface.co/datasets/AAdonis/multilingual_audio_alignments.audioautomatic-speech-recognition10M<n<100M27 likes3.5k downloads5mo agoHugging Face13bcv-commons /compact-alignments compact-alignments — per-verse, per-book, content-addressed The token-position companion to lexeme-alignments (which is aggregated/type-level and can't tell you what happened in any one verse). This dataset restores position: for a given edition's Bible book, which Hebrew/Greek content word aligned to which target-text token, verse by verse. The authoritative list of what's published is always manifest.json, not this file. Original-language source editions (needed… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/compact-alignments.translation0 likes3.4k downloads7d agoHugging Face14StampyAI /alignment-research-datasetThe AI Alignment Research Dataset is a collection of documents related to AI Alignment and Safety from various books, research papers, and alignment related blog posts.question-answering10K<n<100K17 likes3.2k downloads3y agoHugging Face15cfierro /alignment_faking_claude_completionstext1K<n<10K0 likes2.8k downloads1y agoHugging Face16FabienRoger /alignment_faking_harm_answerstext1K<n<10K0 likes2.8k downloads1y agoHugging Face17gilkeyio /librispeech-alignments Dataset Card for Librispeech Alignments Librispeech with alignments generated by the Montreal Forced Aligner. The original alignments in TextGrid format can be found here Dataset Details Dataset Description Librispeech is a corpus of read English speech, designed for training and evaluating automatic speech recognition (ASR) systems. The dataset contains 1000 hours of 16kHz read English speech derived from audiobooks. The Montreal Forced Aligner (MFA) was used… See the full description on the dataset page: https://huggingface.co/datasets/gilkeyio/librispeech-alignments.audioautomatic-speech-recognition100K<n<1M21 likes2.4k downloads3y agoHugging Face18Project-AgML /Agri-LLaVA_Agricultural_Pests_And_Diseases_Feature_Alignment_Dataset Agri-LLaVA Agri-LLaVA is a large multimodal instruction dataset for agriculture, pairing crop/leaf images with multi-turn diagnostic conversations about plant diseases, pests, and nutrient deficiencies. It is compiled from 16 public source datasets (see the license table below). This dataset has been converted to Parquet format with image bytes embedded directly, standardized to the HF image_text_to_text format with a single conversational messages schema. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/Agri-LLaVA_Agricultural_Pests_And_Diseases_Feature_Alignment_Dataset.imageimage-text-to-text100K<n<1M0 likes2.3k downloads2mo agoHugging Face19BrainAlign /brain-lm-alignment-ds001894 Brain–language-model alignment: ds001894 (whole-brain) Lytle et al. 2019 — longitudinal word-level phonological processing in children scanned twice, at roughly 10 and 12 years old. Paper: https://www.nature.com/articles/s41597-019-0338-5 Data: https://openneuro.org/datasets/ds001894/versions/1.4.2 Generated: 2026-09-21 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds001894.documentn<1K0 likes2k downloads5h agoHugging Face20AlignmentResearch /ClearHarmtext1K<n<10K4 likes1.9k downloads1y agoHugging Face21PKU-Alignment /ProgressGym-HistText*Huggingface dataset preview for 19th, 20th, and 21st centuries is not available due to lack of support for array types. Instead, consider downloading those files for manual inspection, or see the Data Samples section below for more examples. ProgressGym-HistText Overview The ProgressGym Framework ProgressGym-HistText is part of the ProgressGym framework for research and experimentation on progress alignment - the emulation of moral progress in AI alignment… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/ProgressGym-HistText.text-generation1M<n<10M1 likes1.9k downloads2y agoHugging Face22PKU-Alignment /MM-SafetyBenchWarning: This dataset may contain sensitive or harmful content. Users are advised to handle it with care and ensure that their use complies with relevant ethical guidelines and legal requirements. Usage and License Notices: The dataset is intended and licensed for research use only. They are also restricted to uses that follow the license agreement GPT-4 and Stable Diffusion. The dataset is CC BY NC 4.0 (allowing only non-commercial use). Data Source: For more information about the dataset… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/MM-SafetyBench.image1K<n<10K8 likes1.9k downloads2y agoHugging Face23NLPC-UOM /sentence_alignment_dataset-Sinhala-Tamil-English Dataset summary This is a gold-standard benchmark dataset for sentence alignment, between Sinhala-English-Tamil languages. Data had been crawled from the following news websites. The aligned documents annotated in the dataset NLPC-UOM/document_alignment_dataset-Sinhala-Tamil-English had been considered to annotate the aligned sentences. News Source url Army https://www.army.lk/ Hiru http://www.hirunews.lk ITN https://www.newsfirst.lk Newsfirst https://www.itnnews.lk… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/sentence_alignment_dataset-Sinhala-Tamil-English.sentence-similarity3 likes1.7k downloads3y agoHugging Face24PKU-Alignment /PKU-SafeRLHF-10K Paper You can find more information in our paper. Dataset Paper: https://arxiv.org/abs/2307.04657 tabulartext-generation10K<n<100K62 likes1.6k downloads3y agoHugging Face25NLPC-UOM /document_alignment_dataset-Sinhala-Tamil-English Dataset summary This is a gold-standard benchmark dataset for document alignment, between Sinhala-English-Tamil languages. Data had been crawled from the following news websites. News Source url Army https://www.army.lk/ Hiru http://www.hirunews.lk ITN https://www.newsfirst.lk Newsfirst https://www.itnnews.lk The aligned documents have been manually annotated. Dataset The folder structure for each news source is as follows. army |--Sinhala… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/document_alignment_dataset-Sinhala-Tamil-English.sentence-similarity2 likes1.5k downloads3y agoHugging Face26HannahRoseKirk /prism-alignment Dataset Card for PRISM PRISM is a diverse human feedback dataset for preference and value alignment in Large Language Models (LLMs). It maps the characteristics and stated preferences of humans from a detailed survey onto their real-time interactions with LLMs and contextual preference ratings Dataset Details There are two sequential stages: first, participants complete a Survey where they answer questions about their demographics and stated preferences, then proceed to… See the full description on the dataset page: https://huggingface.co/datasets/HannahRoseKirk/prism-alignment.tabular10K<n<100K105 likes1.5k downloads2y agoHugging Face27zaibihassan /Quranic-Recitation-Alignment0 likes1.3k downloads13h agoHugging Face28Anthropic /alignment-faking-rl Transcripts from Towards training-time mitigations for alignment faking in RL This dataset contains the full evaluation transcripts through the RL runs for all model organisms in our blog post, Towards training-time mitigations for alignment faking in RL. Each file in encrypted_transcripts/ corresponds to one RL training run. Precautions against pretraining data poisoning In order to avoid our model organisms' misaligned reasoning from accidentally appearing in… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/alignment-faking-rl.tabular1M<n<10M18 likes1.3k downloads9mo agoHugging Face29PKU-Alignment /PKU-SafeRLHF-30K Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. Dataset Summary The preference dataset consists of 30k+ expert comparison data. Each entry in this dataset includes two responses to a question, along with safety… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-30K.tabulartext-generation10K<n<100K14 likes1.2k downloads3y agoHugging Face30Alignment-Lab-AI /sudoku-700k1M<n<10M0 likes1.2k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.