CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01asahi417 /seamless-align-enA-viA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes20k downloads2y agoHugging Face02asahi417 /seamless-align-enA-esA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes20k downloads2y agoHugging Face03asahi417 /seamless-align-enA-frA.speaker-embedding.hubert-xltabular1M<n<10M0 likes16k downloads2y agoHugging Face04PKU-Alignment /BeaverTails Dataset Card for BeaverTails BeaverTails is an AI safety-focused collection comprising a series of datasets. This repository includes human-labeled data consisting of question-answer (QA) pairs, each identified with their corresponding harm categories. It should be noted that a single QA pair can be associated with more than one category. The 14 harm categories are defined as follows: Animal Abuse: This involves any form of cruelty or harm inflicted on animals, including physical… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/BeaverTails.texttext-classification100K<n<1M114 likes16k downloads3y agoHugging Face05asahi417 /seamless-align-enA-jaA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes15k downloads2y agoHugging Face06PKU-Alignment /PKU-SafeRLHF Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. [🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset] Citation If PKU-SafeRLHF has contributed to your work, please consider citing… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF.tabulartext-generation100K<n<1M196 likes15k downloads2y agoHugging Face07asahi417 /seamless-align-deA-enA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes11k downloads2y agoHugging Face08asahi417 /seamless-align-enA-hiA.speaker-embedding.hubert-xltabular100K<n<1M0 likes10k downloads2y agoHugging Face09asahi417 /seamless-align-enA-frA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes9.9k downloads2y agoHugging Face10asahi417 /seamless-align-enA-esA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes9.9k downloads2y agoHugging Face11asahi417 /seamless-align-enA-zhA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes9.7k downloads2y agoHugging Face12asahi417 /seamless-align-enA-frA.speaker-embedding.w2vbert-600mtabular1M<n<10M0 likes8.9k downloads2y agoHugging Face13nguyenvulebinh /asr-alignment Speech Recognition Alignment Dataset This dataset is a variation of several widely-used ASR datasets, encompassing Librispeech, MuST-C, TED-LIUM, VoxPopuli, Common Voice, and GigaSpeech. The difference is this dataset includes: Precise alignment between audio and text. Text that has been punctuated and made case-sensitive. Identification of named entities in the text. Usage First, install the latest version of the 🤗 Datasets package: pip install --upgrade pip pip… See the full description on the dataset page: https://huggingface.co/datasets/nguyenvulebinh/asr-alignment.audio10M<n<100M5 likes8.4k downloads3y agoHugging Face14asahi417 /seamless-align-enA-zhA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes8.1k downloads2y agoHugging Face15asahi417 /seamless-align-enA-zhA.speaker-embedding.hubert-xltabular100K<n<1M0 likes7k downloads2y agoHugging Face16asahi417 /seamless-align-enA-koA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes6.9k downloads2y agoHugging Face17nyu-visionx /Cambrian-Alignment Cambrian-Alignment Dataset Please see paper & website for more information: https://cambrian-mllm.github.io/ https://arxiv.org/abs/2406.16860 Overview Cambrian-Alignment is an question-answering alignment dataset comprised of alignment data from LLaVA, Mini-Gemini, Allava, and ShareGPT4V. Getting Started with Cambrian Alignment Data Before you start, ensure you have sufficient storage space to download and process the data. Download the Data Repository… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-Alignment.imagevisual-question-answering100K<n<1M38 likes6.8k downloads2y agoHugging Face18asahi417 /seamless-align-enA-hiA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes6.5k downloads2y agoHugging Face19dsfsi /govza-sa-cabinet-statements-sentence-aligned Gov-ZA Multilingual Cabinet Statements (Sentence-Aligned) Dataset Description This dataset contains sentence-aligned parallel text from South African government cabinet statements in 11 official languages. The data is sourced from the Government Communication and Information System (GCIS) and scraped from www.gov.za/cabinet-statements. Key Features: 📊 55 language pair combinations covering 11 South African languages 🔗 Sentence-level alignment using LASER embeddings 📈… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi/govza-sa-cabinet-statements-sentence-aligned.texttranslation100K<n<1M1 likes6.3k downloads9mo agoHugging Face20asahi417 /seamless-align-enA-viA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes6.3k downloads2y agoHugging Face21PKU-Alignment /align-anything Overview: Align-Anything Dataset A Comprehensive All-Modality Alignment Dataset with Fine-grained Preference Annotations and Language Feedback. 🏠 Homepage | 🤗 Align-Anything Dataset | 🤗 T2T_Instruction-tuning Dataset | 🤗 TI2T_Instruction-tuning Dataset | 👍 Our Official Code Repo Our world is inherently multimodal. Humans perceive the world through multiple senses, and Language Models should operate similarly. However, the development of Current Multi-Modality Foundation Models… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/align-anything.audioany-to-any10K<n<100K48 likes6.2k downloads1y agoHugging Face22takuM23 /multilingual_audio_alignments Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT) A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.audioautomatic-speech-recognition10M<n<100M4 likes6.1k downloads6mo agoHugging Face23asahi417 /seamless-align-enA-jaA.speaker-embedding.hubert-xltabular100K<n<1M0 likes5.9k downloads2y agoHugging Face24asahi417 /seamless-align-enA-hiA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes5.7k downloads2y agoHugging Face25asahi417 /seamless-align-enA-jaA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes5.6k downloads2y agoHugging Face26AlignmentResearch /DolusChattext10K<n<100K6 likes4.8k downloads1y agoHugging Face27AlignmentResearch /soft-trigger-verifiedtext1K<n<10K0 likes4.5k downloads9mo agoHugging Face28Magpie-Align /Magpie-Pro-300K-Filtered Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Pro-300K-Filtered.text100K<n<1M56 likes4.3k downloads2y agoHugging Face29BrainAlign /brain-lm-alignment-ds002236 Brain–language-model alignment: ds002236 (whole-brain) Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children (8.7–15.5), auditory and visual. Paper: https://pubmed.ncbi.nlm.nih.gov/31956678/ Data: https://openneuro.org/datasets/ds002236/versions/1.0.1 Generated: 2026-09-22 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number in… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds002236.documentn<1K0 likes4.2k downloads27m agoHugging Face30asahi417 /seamless-align-enA-koA.speaker-embedding.hubert-xltabular100K<n<1M0 likes4k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.