alpaca-instruct
gemma-2-9b-alpaca-small-Instruct-GGUFPhi-3-mini-3.8b-4Bit-InstructionTuned-Alpacajpacifico-French-Alpaca-Llama3-8B-Instruct-v1.0-GGUFFireball-Alpaca-Llama-3.1-8B-Instruct-KTO-beta-GGUFjpacifico_-_French-Alpaca-Phi-3-mini-128k-instruct-beta-ggufBanglaLLama-3.1-8b-bangla-alpaca-orca-instruct-v0.0.1Nutanix_-_open-instruct-stanford-alpaca-7b_checkpoint-8850_20241003-065031-merged-ggufSpooke_-_distilgpt2-finetuned-python_code_instructions_18k_alpaca-gguf
python_code_instructions_18k_alpaca
Dataset Card for python_code_instructions_18k_alpaca
The dataset contains problem descriptions and code in python language.
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here.
code_instructions_122k_alpaca_stylecode_instructions_120k_alpaca
Dataset Card for code_instructions_120k_alpaca
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the original source here.
open-instruct-uncensored-alpacaOriginal dataset page from ehartford.
810,102 entries. Sourced from open-instruct-uncensored.jsonl.
Converted the jsonl to a json which can be loaded into something like LLaMa-LoRA-Tuner.
I've also included smaller datasets that includes less entries depending on how much memory you have to work with.
Each one is randomized before being converted, so each dataset is unique in order.
Count of each Dataset:
code_alpaca: 19991
unnatural_instructions: 68231
baize: 166096
self_instruct: 81512… See the full description on the dataset page: https://huggingface.co/datasets/xzuyn/open-instruct-uncensored-alpaca.ru_turbo_alpaca_evol_instructWizardLM_alpaca_evol_instruct_70k_unfilteredThis dataset is the WizardLM dataset victor123/evol_instruct_70k, removing instances of blatant alignment.
54974 instructions remain.
inspired by https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered
All credit to anon8231489123 for the cleanup script that I adapted to wizardlm_clean.py
license: apache-2.0
language:
- en
pretty_name: wizardlm-unfiltered
