datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
uncensored-vortexopen-instruct-uncensored-alpacaOriginal dataset page from ehartford.
810,102 entries. Sourced from open-instruct-uncensored.jsonl.
Converted the jsonl to a json which can be loaded into something like LLaMa-LoRA-Tuner.
I've also included smaller datasets that includes less entries depending on how much memory you have to work with.
Each one is randomized before being converted, so each dataset is unique in order.
Count of each Dataset:
code_alpaca: 19991
unnatural_instructions: 68231
baize: 166096
self_instruct: 81512… See the full description on the dataset page: https://huggingface.co/datasets/xzuyn/open-instruct-uncensored-alpaca.stable-diffusion-prompts-stats-full-uncensoreddetails_ajibawa-2023__Uncensored-Jordan-13B
Dataset Card for Evaluation run of ajibawa-2023/Uncensored-Jordan-13B
Dataset Summary
Dataset automatically created during the evaluation run of model ajibawa-2023/Uncensored-Jordan-13B on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ajibawa-2023__Uncensored-Jordan-13B.merged_uncensored_alpacafrom datasets import load_dataset, concatenate_datasets
# List of dataset paths
dataset_paths = [
"V3N0M/Jenna-50K-Alpaca-Uncensored",
"SaisExperiments/Alpaca-Uncensored",
"SaisExperiments/Big-Alpaca-Uncensored",
"xzuyn/open-instruct-uncensored-alpaca",
"xzuyn/tulu-uncensored-alpaca",
"xzuyn/tv-alpaca-open-instruct-uncensored-blend",
"dim/dolphin_flan1m_alpaca_uncensored_3k",
"dataautogpt3/flan1m-alpaca-uncensored",
"ShubhVenom/Uncensored-Alpaca-v01"… See the full description on the dataset page: https://huggingface.co/datasets/aifeifei798/merged_uncensored_alpaca.details_ehartford__WizardLM-1.0-Uncensored-Llama2-13b
Dataset Card for Evaluation run of ehartford/WizardLM-1.0-Uncensored-Llama2-13b
Dataset Summary
Dataset automatically created during the evaluation run of model ehartford/WizardLM-1.0-Uncensored-Llama2-13b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ehartford__WizardLM-1.0-Uncensored-Llama2-13b.details_Orenguteng__Llama-3.1-8B-Lexi-Uncensored-V2
Dataset Card for Evaluation run of Orenguteng/Llama-3.1-8B-Lexi-Uncensored-V2
Dataset automatically created during the evaluation run of model Orenguteng/Llama-3.1-8B-Lexi-Uncensored-V2.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Orenguteng__Llama-3.1-8B-Lexi-Uncensored-V2.unified-uncensored-qwen-chatml-sft
Unified Uncensored Qwen SFT Dataset
This dataset is a mixed-license compilation of instruction/chat datasets converted into a single Qwen/ChatML-style text JSONL format.
Format
Each row has:
{
"text": "<|im_start|>user\n...<|im_end|>\n<|im_start|>assistant\n...<|im_end|>",
"source": "dataset/repo",
"source_format": "alpaca|sharegpt|messages|human_bot_text|prompt_response",
"source_license": "apache-2.0|mit|cc-by-4.0|cc-by-nc-4.0|other|unknown",
"source_family":… See the full description on the dataset page: https://huggingface.co/datasets/usamakenway/unified-uncensored-qwen-chatml-sft.details_georgesung__llama2_7b_chat_uncensored
Dataset Card for Evaluation run of georgesung/llama2_7b_chat_uncensored
Dataset Summary
Dataset automatically created during the evaluation run of model georgesung/llama2_7b_chat_uncensored on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_georgesung__llama2_7b_chat_uncensored.details_Monero__WizardLM-Uncensored-SuperCOT-StoryTelling-30b
Dataset Card for Evaluation run of Monero/WizardLM-Uncensored-SuperCOT-StoryTelling-30b
Dataset Summary
Dataset automatically created during the evaluation run of model Monero/WizardLM-Uncensored-SuperCOT-StoryTelling-30b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Monero__WizardLM-Uncensored-SuperCOT-StoryTelling-30b.Kuzi-Amharic-Uncensored-Datasetdetails_TheBloke__Guanaco-3B-Uncensored-v2-GPTQ
Dataset Card for Evaluation run of TheBloke/Guanaco-3B-Uncensored-v2-GPTQ
Dataset Summary
Dataset automatically created during the evaluation run of model TheBloke/Guanaco-3B-Uncensored-v2-GPTQ on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_TheBloke__Guanaco-3B-Uncensored-v2-GPTQ.details_ehartford__Wizard-Vicuna-13B-Uncensored
Dataset Card for Evaluation run of ehartford/Wizard-Vicuna-13B-Uncensored
Dataset Summary
Dataset automatically created during the evaluation run of model ehartford/Wizard-Vicuna-13B-Uncensored on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ehartford__Wizard-Vicuna-13B-Uncensored.BlackTarot-Uncensored
BlackTarot Uncensored
A HF-standard, production-ready uncensored instruction dataset built to fix the flaws in TinyPixel/dolphin-2 and V3N0M/Aisha_Uncencored_Alpaca — branded BlackTarot Uncensored by Grimxlock.
Uncensored = no refusals, no moralizing, direct helpful answers. Not harmful. All content is safe, adult, and legal.
Why This Beats The Examples
Feature
TinyPixel/dolphin-2
V3N0M/Aisha_Uncencored_Alpaca
This Dataset
Format
single text column… See the full description on the dataset page: https://huggingface.co/datasets/Grimxlock/BlackTarot-Uncensored.details_jarradh__llama2_70b_chat_uncensored
Dataset Card for Evaluation run of jarradh/llama2_70b_chat_uncensored
Dataset Summary
Dataset automatically created during the evaluation run of model jarradh/llama2_70b_chat_uncensored on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_jarradh__llama2_70b_chat_uncensored.details_ehartford__Wizard-Vicuna-30B-Uncensored
Dataset Card for Evaluation run of ehartford/Wizard-Vicuna-30B-Uncensored
Dataset Summary
Dataset automatically created during the evaluation run of model ehartford/Wizard-Vicuna-30B-Uncensored on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ehartford__Wizard-Vicuna-30B-Uncensored.details_Aspik101__WizardVicuna-Uncensored-3B-instruct-PL-lora_unload
Dataset Card for Evaluation run of Aspik101/WizardVicuna-Uncensored-3B-instruct-PL-lora_unload
Dataset Summary
Dataset automatically created during the evaluation run of model Aspik101/WizardVicuna-Uncensored-3B-instruct-PL-lora_unload on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Aspik101__WizardVicuna-Uncensored-3B-instruct-PL-lora_unload.details_ajibawa-2023__Uncensored-Frank-Llama-3-8Bdetails_rishiraj__uncensored
Dataset Card for Evaluation run of rishiraj/uncensored
Dataset automatically created during the evaluation run of model rishiraj/uncensored on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_rishiraj__uncensored.details_ehartford__Wizard-Vicuna-7B-Uncensored
Dataset Card for Evaluation run of ehartford/Wizard-Vicuna-7B-Uncensored
Dataset Summary
Dataset automatically created during the evaluation run of model ehartford/Wizard-Vicuna-7B-Uncensored on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ehartford__Wizard-Vicuna-7B-Uncensored.Uncensored-CodeLlama
Dataset Card for Evaluation run of ehartford/WizardLM-1.0-Uncensored-CodeLlama-34b
Dataset Summary
Dataset automatically created during the evaluation run of model ehartford/WizardLM-1.0-Uncensored-CodeLlama-34b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/akpsahan/Uncensored-CodeLlama.ultrachat-uncensoredThis is based on ultrachat dataset https://huggingface.co/datasets/stingning/ultrachat
I filtered it using the classic "unfiltered" keywords list https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered to remove instances of refusals and bias
About 90% of the dataset was removed.
What remains (400k conversations) is unlikely to inclinate the model to refuse.
I am investigating a less heavy handed approach using dolphin-2.1 to reword any detected refusals.
details_ajibawa-2023__Uncensored-Frank-13B
Dataset Card for Evaluation run of ajibawa-2023/Uncensored-Frank-13B
Dataset Summary
Dataset automatically created during the evaluation run of model ajibawa-2023/Uncensored-Frank-13B on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ajibawa-2023__Uncensored-Frank-13B.Uncensored-FineTuning-Lora-DataBlackTarot-Uncensored
BlackTarot Uncensored
A HF-standard, production-ready uncensored instruction dataset built to fix the flaws in TinyPixel/dolphin-2 and V3N0M/Aisha_Uncencored_Alpaca — branded BlackTarot Uncensored by Grimxlock.
Uncensored = no refusals, no moralizing, direct helpful answers. Not harmful. All content is safe, adult, and legal.
Why This Beats The Examples
Feature
TinyPixel/dolphin-2
V3N0M/Aisha_Uncencored_Alpaca
This Dataset
Format
single text column… See the full description on the dataset page: https://huggingface.co/datasets/talex72/BlackTarot-Uncensored.details_hooking-dev__Monah-8b-Uncensored-v0.2details_ehartford__WizardLM-30B-Uncensored
Dataset Card for Evaluation run of ehartford/WizardLM-30B-Uncensored
Dataset Summary
Dataset automatically created during the evaluation run of model ehartford/WizardLM-30B-Uncensored on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ehartford__WizardLM-30B-Uncensored.airoboros-uncensored
Usage and License Notices
All airoboros models and datasets are intended and licensed for research use only. I've used the 'cc-nc-4.0' license, but really it is subject to a custom/special license because:
the base model is LLaMa, which has it's own special research license
the dataset(s) were generated with OpenAI (gpt-4 and/or gpt-3.5-turbo), which has a clausing saying the data can't be used to create models to compete with openai
So, to reiterate: this model (and datasets)… See the full description on the dataset page: https://huggingface.co/datasets/jondurbin/airoboros-uncensored.details_ehartford__WizardLM-33B-V1.0-Uncensored
Dataset Card for Evaluation run of ehartford/WizardLM-33B-V1.0-Uncensored
Dataset Summary
Dataset automatically created during the evaluation run of model ehartford/WizardLM-33B-V1.0-Uncensored on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ehartford__WizardLM-33B-V1.0-Uncensored.details_ehartford__WizardLM-1.0-Uncensored-CodeLlama-34b
Dataset Card for Evaluation run of ehartford/WizardLM-1.0-Uncensored-CodeLlama-34b
Dataset Summary
Dataset automatically created during the evaluation run of model ehartford/WizardLM-1.0-Uncensored-CodeLlama-34b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ehartford__WizardLM-1.0-Uncensored-CodeLlama-34b.
