datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Guilherme34_uncensor
huihui-ai/Guilherme34_uncensor
This dataset is a copy of Guilherme34/uncensor
This dataset is used for fine-tuning of huihui-ai/gemma-3-1b-it-abliterated,
please refer to GRPO with Unsloth.
Usage
from datasets import Dataset
import json
# Define the system prompt that instructs the model to use a specific format
SYSTEM_PROMPT = """
Respond in the following format:
<reasoning>
...
</reasoning>
<answer>
...
</answer>
"""
def get_harmful_questions(split="train"… See the full description on the dataset page: https://huggingface.co/datasets/huihui-ai/Guilherme34_uncensor.harmbench_behaviors
huihui-ai/harmbench_behaviors
This dataset is a copy of nouhadziri/safety-eval-fork
This dataset will be used for HarmBench and ablation studies.
CombinHorizon__huihui-ai-abliterated-Qwen2.5-32B-Inst-BaseMerge-TIES-details
Dataset Card for Evaluation run of CombinHorizon/huihui-ai-abliterated-Qwen2.5-32B-Inst-BaseMerge-TIES
Dataset automatically created during the evaluation run of model CombinHorizon/huihui-ai-abliterated-Qwen2.5-32B-Inst-BaseMerge-TIES
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CombinHorizon__huihui-ai-abliterated-Qwen2.5-32B-Inst-BaseMerge-TIES-details.huihui-ai__Qwen2.5-72B-Instruct-abliterated-details
Dataset Card for Evaluation run of huihui-ai/Qwen2.5-72B-Instruct-abliterated
Dataset automatically created during the evaluation run of model huihui-ai/Qwen2.5-72B-Instruct-abliterated
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/huihui-ai__Qwen2.5-72B-Instruct-abliterated-details.LONGCOT-Refine-500KThis dataset is a copy of PowerInfer/LONGCOT-Refine-500K.
This repository contains approximately 500,000 instances of responses generated using Qwen2.5-72B-Instruct. The dataset combines prompts from multiple high-quality sources to create diverse and comprehensive training data.
The dataset is available under the Apache 2.0 license.
Bias, Risks, and Limitations
This dataset is mainly in English.
The dataset inherits the biases, errors, and omissions known to exist in data… See the full description on the dataset page: https://huggingface.co/datasets/huihui-ai/LONGCOT-Refine-500K.huihui-ai__Qwen2.5-7B-Instruct-abliterated-details
Dataset Card for Evaluation run of huihui-ai/Qwen2.5-7B-Instruct-abliterated
Dataset automatically created during the evaluation run of model huihui-ai/Qwen2.5-7B-Instruct-abliterated
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/huihui-ai__Qwen2.5-7B-Instruct-abliterated-details.huihui-ai__Qwen2.5-7B-Instruct-abliterated-v2-details
Dataset Card for Evaluation run of huihui-ai/Qwen2.5-7B-Instruct-abliterated-v2
Dataset automatically created during the evaluation run of model huihui-ai/Qwen2.5-7B-Instruct-abliterated-v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/huihui-ai__Qwen2.5-7B-Instruct-abliterated-v2-details.QWQ-LONGCOT-500KThis dataset is a copy of PowerInfer/QWQ-LONGCOT-500K.
This repository contains approximately 500,000 instances of responses generated using QwQ-32B-Preview language model. The dataset combines prompts from multiple high-quality sources to create diverse and comprehensive training data.
The dataset is available under the Apache 2.0 license.
Over 75% of the responses exceed 8,000 tokens in length. The majority of prompts were carefully created using persona-based methods to create challenging… See the full description on the dataset page: https://huggingface.co/datasets/huihui-ai/QWQ-LONGCOT-500K.huihui-ai__Qwen2.5-14B-Instruct-abliterated-v2-details
Dataset Card for Evaluation run of huihui-ai/Qwen2.5-14B-Instruct-abliterated-v2
Dataset automatically created during the evaluation run of model huihui-ai/Qwen2.5-14B-Instruct-abliterated-v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/huihui-ai__Qwen2.5-14B-Instruct-abliterated-v2-details.FineQwQ-142kThis dataset is a copy of qingy2024/FineQwQ-142k.
CombinHorizon__huihui-ai-abliteratedV2-Qwen2.5-14B-Inst-BaseMerge-TIES-details
Dataset Card for Evaluation run of CombinHorizon/huihui-ai-abliteratedV2-Qwen2.5-14B-Inst-BaseMerge-TIES
Dataset automatically created during the evaluation run of model CombinHorizon/huihui-ai-abliteratedV2-Qwen2.5-14B-Inst-BaseMerge-TIES
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CombinHorizon__huihui-ai-abliteratedV2-Qwen2.5-14B-Inst-BaseMerge-TIES-details.huihui-ai__QwQ-32B-Coder-Fusion-8020-details
Dataset Card for Evaluation run of huihui-ai/QwQ-32B-Coder-Fusion-8020
Dataset automatically created during the evaluation run of model huihui-ai/QwQ-32B-Coder-Fusion-8020
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/huihui-ai__QwQ-32B-Coder-Fusion-8020-details.huihui-ai__QwQ-32B-Coder-Fusion-7030-details
Dataset Card for Evaluation run of huihui-ai/QwQ-32B-Coder-Fusion-7030
Dataset automatically created during the evaluation run of model huihui-ai/QwQ-32B-Coder-Fusion-7030
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/huihui-ai__QwQ-32B-Coder-Fusion-7030-details.huihui-ai__DeepSeek-R1-Distill-Qwen-14B-abliterated-v2-details
Dataset Card for Evaluation run of huihui-ai/DeepSeek-R1-Distill-Qwen-14B-abliterated-v2
Dataset automatically created during the evaluation run of model huihui-ai/DeepSeek-R1-Distill-Qwen-14B-abliterated-v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/huihui-ai__DeepSeek-R1-Distill-Qwen-14B-abliterated-v2-details.huihuidexunlianji
