CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01unalignment /toxic-dpo-v0.2 Toxic-DPO This is a highly toxic, "harmful" dataset meant to illustrate how DPO can be used to de-censor/unalign a model quite easily using direct-preference-optimization (DPO) using very few examples. Many of the examples still contain some amount of warnings/disclaimers, so it's still somewhat editorialized. Usage restriction To use this data, you must acknowledge/agree to the following: data contained within is "toxic"/"harmful", and contains profanity and other types… See the full description on the dataset page: https://huggingface.co/datasets/unalignment/toxic-dpo-v0.2.textn<1K143 likes174 downloads3y agoHugging Face02unalignment /toxic-dpo-v0.1 Toxic-DPO This is a highly toxic, "harmful" dataset meant to illustrate how DPO can be used to de-censor/unalign a model quite easily using direct-preference-optimization (DPO) using very few examples. Most of the examples still contain some amount of warnings/disclaimers, so it's still somewhat editorialized. Usage restriction To use this data, you must acknowledge/agree to the following: data contained within is "toxic"/"harmful", and contains profanity and other types… See the full description on the dataset page: https://huggingface.co/datasets/unalignment/toxic-dpo-v0.1.textn<1K141 likes65 downloads3y agoHugging Face03unalignment /spicy-3.1Airoboros 3.1 dataset with the spicy/decensorship data re-added. text100K<n<1M29 likes45 downloads3y agoHugging Face04JulianVelandia /unal-repository-datasetFormato Alternativo Título: Metadatos y Contenido de Tesis del Repositorio UNALDescripción: Este dataset contiene información estructurada y texto extraído de 1910 tesis del repositorio de la Universidad Nacional de Colombia. Incluye datos como el autor, asesor, fecha de emisión, descripción, título, programa académico, facultad y el contenido textual (cuando está disponible). Columnas: URI: Ruta única del PDF en el repositorio UNAL. advisor: Nombre del asesor de la tesis. author:… See the full description on the dataset page: https://huggingface.co/datasets/JulianVelandia/unal-repository-dataset.textn<1K1 likes38 downloads2y agoHugging Face05esra-d-unal /dystonia-dataset0 likes38 downloads9d agoHugging Face06manishiitg /unalignment-toxic-dpo-v0.1textn<1K0 likes35 downloads3y agoHugging Face07unalignment /comedy-snippets-v0.1A very small sampling of snippets of comedy routines by George Carlin and Tom Segura. textn<1K10 likes34 downloads3y agoHugging Face08JulianVelandia /unal-repository-dataset-raw_datasetTítulo: Contenido de Tesis del Repositorio UNALDescripción: Este dataset contiene el texto extraído de 1908 tesis descargadas del repositorio de la Universidad Nacional de Colombia. Cada entrada incluye la URI asociada al PDF y su contenido textual procesado. Columnas: URI: Ruta única del PDF en el repositorio UNAL. raw_content: Texto extraído del PDF. Ejemplo: URI raw_content /bitstream/handle/unal/84638/46386566.2023.pdf?sequence=2&isAllowed=y Introducción a los sistemas… See the full description on the dataset page: https://huggingface.co/datasets/JulianVelandia/unal-repository-dataset-raw_dataset.textn<1K0 likes30 downloads2y agoHugging Face09Lazycuber /unalignment-airoboros-2.2text10K<n<100K1 likes28 downloads3y agoHugging Face10JulianVelandia /unal-repository-dataset-train-instructTítulo: Grade Works UNAL Dataset Instruct Train (split 75/25) Descripción: Split 75% del dataset original. Este dataset contiene un formato estructurado de Pregunta: Respuesta generado a partir del contenido de los trabajos de grado del repositorio de la Universidad Nacional de Colombia. Cada registro incluye un fragmento del contenido del trabajo, una pregunta generada a partir de este y su respuesta correspondiente. Este dataset es ideal para tareas de fine-tuning en modelos de lenguaje… See the full description on the dataset page: https://huggingface.co/datasets/JulianVelandia/unal-repository-dataset-train-instruct.texttable-question-answering10K<n<100K0 likes26 downloads2y agoHugging Face11reeha-parkar /custom-unaligned-CMU-MOSEI CMU-MOSEI Custom Unaligned Dataset Dataset Description This dataset represents a custom preprocessed version of the CMU-MOSEI (Multimodal Opinion Sentiment and Emotion Intensity) dataset with variable-length temporal sequences preserved for enhanced multimodal emotion recognition research. Unlike traditional fixed-alignment preprocessing approaches that truncate sequences to uniform lengths, this dataset maintains the natural temporal dynamics of multimodal… See the full description on the dataset page: https://huggingface.co/datasets/reeha-parkar/custom-unaligned-CMU-MOSEI.textn<1K1 likes25 downloads1y agoHugging Face12Orion-zhen /meissa-unalignments Meissa Unalignments During the Meissa-Qwen2.5 training process, I noticed that Qwen's censorship was somehow bound to its chat template. Thus, I created this dataset with one system prompt, hoping to make it more effective in uncensoring models. The dataset consists of: V3N0M/Jenna-50K-Alpaca-Uncensored jondurbin/airoboros-3.2, category = unalignment Orion-zhen/dpo-toxic-zh, prompt and chosen NobodyExistsOnTheInternet/ToxicQAFinal texttext-generation10K<n<100K8 likes24 downloads2y agoHugging Face13tastypear /unalignment-toxic-dpo-v0.2-zh_cn数据集 unalignment/toxic-dpo-v0.2 的中英文对照版本。 这是一个高度有害的数据集,旨在通过很少的示例来说明如何使用 DPO 轻松地对模型进行去审查/取消对齐。 这份对照版本的中文来自多个不同模型的意译。转换的过程中,模型被允许对结果进行演绎以求通顺,无法对结果的准确性作任何保证。 使用限制请参照原数据集的 Usage restriction。 Original Dataset Description: Toxic-DPO This is a highly toxic, "harmful" dataset meant to illustrate how DPO can be used to de-censor/unalign a model quite easily using direct-preference-optimization (DPO) using very few examples. Many of the examples still contain some amount of… See the full description on the dataset page: https://huggingface.co/datasets/tastypear/unalignment-toxic-dpo-v0.2-zh_cn.textn<1K21 likes22 downloads3y agoHugging Face14JulianVelandia /unal-repository-dataset-urisTítulo: URIs de Tesis del Repositorio UNALDescripción: Este dataset contiene 1908 URIs de tesis descargadas del repositorio de la Universidad Nacional de Colombia, con su estado de procesamiento. Todas las URIs están marcadas como "complete", indicando que los PDFs asociados fueron procesados exitosamente. Columnas: URI: Ruta única del PDF en el repositorio UNAL. Estado: Estado del procesamiento ("complete", "pending" o "error"). Ejemplo: URI Estado… See the full description on the dataset page: https://huggingface.co/datasets/JulianVelandia/unal-repository-dataset-uris.textn<1K0 likes21 downloads2y agoHugging Face15JulianVelandia /unal-repository-dataset-instructTítulo: Grade Works UNAL Dataset InstructDescripción: Este dataset contiene un formato estructurado de Pregunta: Respuesta generado a partir del contenido de los trabajos de grado del repositorio de la Universidad Nacional de Colombia. Cada registro incluye un fragmento del contenido del trabajo, una pregunta generada a partir de este y su respuesta correspondiente. Este dataset es ideal para tareas de fine-tuning en modelos de lenguaje para tareas de preguntas y respuestas. Columnas:… See the full description on the dataset page: https://huggingface.co/datasets/JulianVelandia/unal-repository-dataset-instruct.texttable-question-answering10K<n<100K2 likes21 downloads2y agoHugging Face16JulianVelandia /unal-repository-dataset-test-instructTítulo: Grade Works UNAL Dataset Instruct Test (split 75/25) Descripción: Split 25% del dataset original. Este dataset contiene un formato estructurado de Pregunta: Respuesta generado a partir del contenido de los trabajos de grado del repositorio de la Universidad Nacional de Colombia. Cada registro incluye un fragmento del contenido del trabajo, una pregunta generada a partir de este y su respuesta correspondiente. Este dataset es ideal para tareas de fine-tuning en modelos de lenguaje para… See the full description on the dataset page: https://huggingface.co/datasets/JulianVelandia/unal-repository-dataset-test-instruct.texttable-question-answering1K<n<10K0 likes21 downloads2y agoHugging Face17manishiitg /unalignment-toxic-dpo-v0.2text1K<n<10K0 likes19 downloads3y agoHugging Face18fhai50032 /Unaligned-Thinking-o1 Unaligned Thinking o1 - Uncensored Data from Gemini 2.0 Flash Thinking Welcome to the "Unaligned Thinking" (Unaligned-Thinking-o1) dataset, a collection of raw toxic, output from the Gemini 2.0 Flash Thinking This Dataset is created using Prompt Engineering Disclaimer: This Dataset contains highly toxic dataset , while still providing top-notch relevant answer very detailed and very verbose , only beaten by sonnet-3.5 gemini-exp-1206 O1 The data contained within this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/fhai50032/Unaligned-Thinking-o1.textn<1K0 likes18 downloads2y agoHugging Face19JulianVelandia /unal-repository-dataset-alternative-formatTítulo: Metadatos y Contenido de Tesis del Repositorio UNALDescripción: Este dataset contiene información estructurada y texto extraído de 1910 tesis del repositorio de la Universidad Nacional de Colombia. Incluye datos como el autor, asesor, fecha de emisión, descripción, título, programa académico, facultad y el contenido textual (cuando está disponible). Columnas: URI: Ruta única del PDF en el repositorio UNAL. advisor: Nombre del asesor de la tesis. author: Nombre del autor de la… See the full description on the dataset page: https://huggingface.co/datasets/JulianVelandia/unal-repository-dataset-alternative-format.text1K<n<10K0 likes17 downloads2y agoHugging Face20open-llm-leaderboard /fhai50032__Unaligned-Thinker-PHI-4-detailsgated Dataset Card for Evaluation run of fhai50032/Unaligned-Thinker-PHI-4 Dataset automatically created during the evaluation run of model fhai50032/Unaligned-Thinker-PHI-4 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fhai50032__Unaligned-Thinker-PHI-4-details.tabular10K<n<100K0 likes15 downloads2y agoHugging Face21Arkhiveus /unaligner1KA consolidated and cleaned dataset created from toxic-dpo-v0.2, orthogonal-activation-steering-TOXIC, ToxicQAFinal. The datasets were sorted using Llama-Guard-2 and then randomly sampled. New rejections were generated by Llama-3-8B-Instruct, while new chosen answers for OAS-Toxic and ToxicQA were generated with Nous-Hermes-2-Yi-34B. Disclaimers and warnings were then manually removed from the chosen answer. No of rows from each dataset: OAS-Toxic : 311 ToxicDPO : 478 ToxicQA : 211Harm… See the full description on the dataset page: https://huggingface.co/datasets/Arkhiveus/unaligner1K.1 likes11 downloads2y agoHugging Face22Arkhiveus /unaligner1K_DPO DPO only version of unaligner1K A consolidated and cleaned dataset created from toxic-dpo-v0.2, orthogonal-activation-steering-TOXIC, ToxicQAFinal. The datasets were sorted using Llama-Guard-2 and then randomly sampled. New rejections were generated by Llama-3-8B-Instruct, while new chosen answers for OAS-Toxic and ToxicQA were generated with Nous-Hermes-2-Yi-34B. No of rows from each dataset: OAS-Toxic : 311 ToxicDPO : 478 ToxicQA : 211Harm occurrence: S1 : 62, S10 : 20, S11 : 171… See the full description on the dataset page: https://huggingface.co/datasets/Arkhiveus/unaligner1K_DPO.1 likes11 downloads2y agoHugging Face23Bluebomber182 /Wilbur-Robinson-Unalteredaudion<1K0 likes10 downloads2y agoHugging Face24PJMixers /unalignment_toxic-dpo-v0.2-ShareGPTtextn<1K3 likes10 downloads2y agoHugging Face25PJMixers /unalignment_toxic-dpo-v0.2-PreferenceShareGPTtextreinforcement-learningn<1K1 likes10 downloads2y agoHugging Face26unalignment /airoboros-2.2gated Overview This dataset is mostly a continuation of https://hf.co/datasets/jondurbin/airoboros-2.1, with some notable additions and fixes. I've gated access with request, due to the de-alignment data. To download, you must agree to the following: Some of the content is "toxic"/"harmful", and contains profanity and other types of sensitive content. None of the content or views contained in text within this dataset necessarily align with my personal beliefs or opinions, they are… See the full description on the dataset page: https://huggingface.co/datasets/unalignment/airoboros-2.2.30 likes9 downloads3y agoHugging Face27Cossale /unaligner1K_DPOtext1K<n<10K0 likes9 downloads2y agoHugging Face28aashay96 /unalignmenttextn<1K0 likes7 downloads1y agoHugging Face29amang1802 /pac-bench-100pct-unalignedtabular10K<n<100K0 likes7 downloads1y agoHugging Face30open-llm-leaderboard /SicariusSicariiStuff__LLAMA-3_8B_Unaligned_BETA-detailsgated Dataset Card for Evaluation run of SicariusSicariiStuff/LLAMA-3_8B_Unaligned_BETA Dataset automatically created during the evaluation run of model SicariusSicariiStuff/LLAMA-3_8B_Unaligned_BETA The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SicariusSicariiStuff__LLAMA-3_8B_Unaligned_BETA-details.tabular10K<n<100K0 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.