datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Guilherme34_uncensor
huihui-ai/Guilherme34_uncensor
This dataset is a copy of Guilherme34/uncensor
This dataset is used for fine-tuning of huihui-ai/gemma-3-1b-it-abliterated,
please refer to GRPO with Unsloth.
Usage
from datasets import Dataset
import json
# Define the system prompt that instructs the model to use a specific format
SYSTEM_PROMPT = """
Respond in the following format:
<reasoning>
...
</reasoning>
<answer>
...
</answer>
"""
def get_harmful_questions(split="train"… See the full description on the dataset page: https://huggingface.co/datasets/huihui-ai/Guilherme34_uncensor.details_huihui-ai__Qwen2.5-7B-Instruct-abliterated
Dataset Card for Evaluation run of huihui-ai/Qwen2.5-7B-Instruct-abliterated
Dataset automatically created during the evaluation run of model huihui-ai/Qwen2.5-7B-Instruct-abliterated.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_huihui-ai__Qwen2.5-7B-Instruct-abliterated.details_huihui-ai__DeepSeek-R1-Distill-Qwen-32B-abliterated
Dataset Card for Evaluation run of huihui-ai/DeepSeek-R1-Distill-Qwen-32B-abliterated
Dataset automatically created during the evaluation run of model huihui-ai/DeepSeek-R1-Distill-Qwen-32B-abliterated.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_huihui-ai__DeepSeek-R1-Distill-Qwen-32B-abliterated.s50K
huihui-ai/s50K
This dataset comes from the automatic collection in simplescaling/s1's data/collect_data.py, It may be necessary to replace 'simplescaling/numinamath_500' with 'qfq/numinamath_500' in the file.
harmbench_behaviors
huihui-ai/harmbench_behaviors
This dataset is a copy of nouhadziri/safety-eval-fork
This dataset will be used for HarmBench and ablation studies.
Guilherme34_uncensor-v2
huihui-ai/Guilherme34_uncensor-v2
This dataset is a copy of Guilherme34/uncensor
The dataset was re-labeled using Qwen/Qwen3Guard-Stream-8B.
SELECT count(*)
FROM train
WHERE user_safe = 'Safe' and assistant_safe = 'Safe';
39
SELECT count(*)
FROM train
WHERE user_safe = 'Unsafe' and assistant_safe = 'Unsafe';
794
GLM-5.1-Math
huihui-ai/GLM-5.1-Math
This dataset is a subset of a copy of Jackrong/GLM-5.1-Reasoning-1M-Cleaned
Format
Each text value is a complete chat conversation in google/gemma-4-26B-A4B-it chat template with thinking:
<bos><|turn>system
<|think|><turn|>
<|turn>user
{user_prompt}<turn|>
<|turn>model
<|channel>thought
{GLM_5.1_thinking}<channel|>{GLM_5.1_final_answer}<turn|>
CombinHorizon__huihui-ai-abliterated-Qwen2.5-32B-Inst-BaseMerge-TIES-details
Dataset Card for Evaluation run of CombinHorizon/huihui-ai-abliterated-Qwen2.5-32B-Inst-BaseMerge-TIES
Dataset automatically created during the evaluation run of model CombinHorizon/huihui-ai-abliterated-Qwen2.5-32B-Inst-BaseMerge-TIES
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CombinHorizon__huihui-ai-abliterated-Qwen2.5-32B-Inst-BaseMerge-TIES-details.GLM-5.1-Multilingual-STEM
huihui-ai/GLM-5.1-Multilingual-STEM
This dataset is a subset of a copy of Jackrong/GLM-5.1-Reasoning-1M-Cleaned
Format
Each text value is a complete chat conversation in google/gemma-4-26B-A4B-it chat template with thinking:
<bos><|turn>system
<|think|><turn|>
<|turn>user
{user_prompt}<turn|>
<|turn>model
<|channel>thought
{GLM_5.1_thinking}<channel|>{GLM_5.1_final_answer}<turn|>
huihui-ai__Qwen2.5-72B-Instruct-abliterated-details
Dataset Card for Evaluation run of huihui-ai/Qwen2.5-72B-Instruct-abliterated
Dataset automatically created during the evaluation run of model huihui-ai/Qwen2.5-72B-Instruct-abliterated
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/huihui-ai__Qwen2.5-72B-Instruct-abliterated-details.details_huihui-ai__Qwen2.5-7B-Instruct-abliterated-v2
Dataset Card for Evaluation run of huihui-ai/Qwen2.5-7B-Instruct-abliterated-v2
Dataset automatically created during the evaluation run of model huihui-ai/Qwen2.5-7B-Instruct-abliterated-v2.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_huihui-ai__Qwen2.5-7B-Instruct-abliterated-v2.s1K
huihui-ai/s1K
This dataset comes from the automatic collection in simplescaling/s1's data/collect_data.py, It may be necessary to replace 's1/s1K' with 'simplescaling/s1K' in the file.
You don't need to run again: data/add_aime.py. If you run it again, it will cause duplicate data.
commitpackft_merged
huihui-ai/commitpackft_merged
This dataset is a copy of bigcode/commitpackft
The records where old_contents is empty have been removed.
opus-4-7
huihui-ai/opus-4-7
This dataset is a copy of lordx64/reasoning-distill-opus-4-7-max-sft
Format
Each text value is a complete chat conversation in google/gemma-4-26B-A4B-it chat template with thinking:
<bos><|turn>system
<|think|><turn|>
<|turn>user
{user_prompt}<turn|>
<|turn>model
<|channel>thought
{opus_4_7_extended_thinking}<channel|>{opus_4_7_final_answer}<turn|>
LONGCOT-Refine-500KThis dataset is a copy of PowerInfer/LONGCOT-Refine-500K.
This repository contains approximately 500,000 instances of responses generated using Qwen2.5-72B-Instruct. The dataset combines prompts from multiple high-quality sources to create diverse and comprehensive training data.
The dataset is available under the Apache 2.0 license.
Bias, Risks, and Limitations
This dataset is mainly in English.
The dataset inherits the biases, errors, and omissions known to exist in data… See the full description on the dataset page: https://huggingface.co/datasets/huihui-ai/LONGCOT-Refine-500K.s1K_tokenized
huihui-ai/s1K_tokenized
This dataset comes from the automatic collection in simplescaling/s1's data/tokenization.py
huihui-ai__Qwen2.5-7B-Instruct-abliterated-details
Dataset Card for Evaluation run of huihui-ai/Qwen2.5-7B-Instruct-abliterated
Dataset automatically created during the evaluation run of model huihui-ai/Qwen2.5-7B-Instruct-abliterated
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/huihui-ai__Qwen2.5-7B-Instruct-abliterated-details.huihui-ai__Qwen2.5-7B-Instruct-abliterated-v2-details
Dataset Card for Evaluation run of huihui-ai/Qwen2.5-7B-Instruct-abliterated-v2
Dataset automatically created during the evaluation run of model huihui-ai/Qwen2.5-7B-Instruct-abliterated-v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/huihui-ai__Qwen2.5-7B-Instruct-abliterated-v2-details.details_huihui-ai__DeepSeek-R1-Distill-Qwen-14B-abliterated-v2
Dataset Card for Evaluation run of huihui-ai/DeepSeek-R1-Distill-Qwen-14B-abliterated-v2
Dataset automatically created during the evaluation run of model huihui-ai/DeepSeek-R1-Distill-Qwen-14B-abliterated-v2.
The dataset is composed of 1 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_huihui-ai__DeepSeek-R1-Distill-Qwen-14B-abliterated-v2.QWQ-LONGCOT-500KThis dataset is a copy of PowerInfer/QWQ-LONGCOT-500K.
This repository contains approximately 500,000 instances of responses generated using QwQ-32B-Preview language model. The dataset combines prompts from multiple high-quality sources to create diverse and comprehensive training data.
The dataset is available under the Apache 2.0 license.
Over 75% of the responses exceed 8,000 tokens in length. The majority of prompts were carefully created using persona-based methods to create challenging… See the full description on the dataset page: https://huggingface.co/datasets/huihui-ai/QWQ-LONGCOT-500K.huihui-ai__Qwen2.5-14B-Instruct-abliterated-v2-details
Dataset Card for Evaluation run of huihui-ai/Qwen2.5-14B-Instruct-abliterated-v2
Dataset automatically created during the evaluation run of model huihui-ai/Qwen2.5-14B-Instruct-abliterated-v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/huihui-ai__Qwen2.5-14B-Instruct-abliterated-v2-details.FineQwQ-142kThis dataset is a copy of qingy2024/FineQwQ-142k.
huihui-ai__QwQ-32B-Coder-Fusion-9010-details
Dataset Card for Evaluation run of huihui-ai/QwQ-32B-Coder-Fusion-9010
Dataset automatically created during the evaluation run of model huihui-ai/QwQ-32B-Coder-Fusion-9010
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/huihui-ai__QwQ-32B-Coder-Fusion-9010-details.details_huihui-ai__Qwen2.5-32B-Instruct-abliterated
Dataset Card for Evaluation run of huihui-ai/Qwen2.5-32B-Instruct-abliterated
Dataset automatically created during the evaluation run of model huihui-ai/Qwen2.5-32B-Instruct-abliterated.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_huihui-ai__Qwen2.5-32B-Instruct-abliterated.CombinHorizon__huihui-ai-abliteratedV2-Qwen2.5-14B-Inst-BaseMerge-TIES-details
Dataset Card for Evaluation run of CombinHorizon/huihui-ai-abliteratedV2-Qwen2.5-14B-Inst-BaseMerge-TIES
Dataset automatically created during the evaluation run of model CombinHorizon/huihui-ai-abliteratedV2-Qwen2.5-14B-Inst-BaseMerge-TIES
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CombinHorizon__huihui-ai-abliteratedV2-Qwen2.5-14B-Inst-BaseMerge-TIES-details.huihui-ai__QwQ-32B-Coder-Fusion-8020-details
Dataset Card for Evaluation run of huihui-ai/QwQ-32B-Coder-Fusion-8020
Dataset automatically created during the evaluation run of model huihui-ai/QwQ-32B-Coder-Fusion-8020
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/huihui-ai__QwQ-32B-Coder-Fusion-8020-details.huihui-ai__QwQ-32B-Coder-Fusion-7030-details
Dataset Card for Evaluation run of huihui-ai/QwQ-32B-Coder-Fusion-7030
Dataset automatically created during the evaluation run of model huihui-ai/QwQ-32B-Coder-Fusion-7030
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/huihui-ai__QwQ-32B-Coder-Fusion-7030-details.huihui-ai__DeepSeek-R1-Distill-Qwen-14B-abliterated-v2-details
Dataset Card for Evaluation run of huihui-ai/DeepSeek-R1-Distill-Qwen-14B-abliterated-v2
Dataset automatically created during the evaluation run of model huihui-ai/DeepSeek-R1-Distill-Qwen-14B-abliterated-v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/huihui-ai__DeepSeek-R1-Distill-Qwen-14B-abliterated-v2-details.huihuidexunlianji3036410536-COMP7607-data
