uncensored
uncensored-vortexopen-instruct-uncensored-alpacaOriginal dataset page from ehartford.
810,102 entries. Sourced from open-instruct-uncensored.jsonl.
Converted the jsonl to a json which can be loaded into something like LLaMa-LoRA-Tuner.
I've also included smaller datasets that includes less entries depending on how much memory you have to work with.
Each one is randomized before being converted, so each dataset is unique in order.
Count of each Dataset:
code_alpaca: 19991
unnatural_instructions: 68231
baize: 166096
self_instruct: 81512… See the full description on the dataset page: https://huggingface.co/datasets/xzuyn/open-instruct-uncensored-alpaca.stable-diffusion-prompts-stats-full-uncensoreddetails_ajibawa-2023__Uncensored-Jordan-13B
Dataset Card for Evaluation run of ajibawa-2023/Uncensored-Jordan-13B
Dataset Summary
Dataset automatically created during the evaluation run of model ajibawa-2023/Uncensored-Jordan-13B on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ajibawa-2023__Uncensored-Jordan-13B.merged_uncensored_alpacafrom datasets import load_dataset, concatenate_datasets
# List of dataset paths
dataset_paths = [
"V3N0M/Jenna-50K-Alpaca-Uncensored",
"SaisExperiments/Alpaca-Uncensored",
"SaisExperiments/Big-Alpaca-Uncensored",
"xzuyn/open-instruct-uncensored-alpaca",
"xzuyn/tulu-uncensored-alpaca",
"xzuyn/tv-alpaca-open-instruct-uncensored-blend",
"dim/dolphin_flan1m_alpaca_uncensored_3k",
"dataautogpt3/flan1m-alpaca-uncensored",
"ShubhVenom/Uncensored-Alpaca-v01"… See the full description on the dataset page: https://huggingface.co/datasets/aifeifei798/merged_uncensored_alpaca.details_ehartford__WizardLM-1.0-Uncensored-Llama2-13b
Dataset Card for Evaluation run of ehartford/WizardLM-1.0-Uncensored-Llama2-13b
Dataset Summary
Dataset automatically created during the evaluation run of model ehartford/WizardLM-1.0-Uncensored-Llama2-13b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ehartford__WizardLM-1.0-Uncensored-Llama2-13b.
