Instruct Alpaca
gemma-2-9b-alpaca-small-Instruct-GGUFPhi-3-mini-3.8b-4Bit-InstructionTuned-Alpacajpacifico-French-Alpaca-Llama3-8B-Instruct-v1.0-GGUFFireball-Alpaca-Llama-3.1-8B-Instruct-KTO-beta-GGUFBanglaLLama-3.1-8b-bangla-alpaca-orca-instruct-v0.0.1jpacifico_-_French-Alpaca-Phi-3-mini-128k-instruct-beta-ggufBanglaLLama-3.2-3b-bangla-alpaca-orca-instruct-v0.0.1-GGUFFrench-Alpaca-Llama3-8B-Instruct-v1.0-GGUF
open-instruct-uncensored-alpacaOriginal dataset page from ehartford.
810,102 entries. Sourced from open-instruct-uncensored.jsonl.
Converted the jsonl to a json which can be loaded into something like LLaMa-LoRA-Tuner.
I've also included smaller datasets that includes less entries depending on how much memory you have to work with.
Each one is randomized before being converted, so each dataset is unique in order.
Count of each Dataset:
code_alpaca: 19991
unnatural_instructions: 68231
baize: 166096
self_instruct: 81512… See the full description on the dataset page: https://huggingface.co/datasets/xzuyn/open-instruct-uncensored-alpaca.WizardLM_alpaca_evol_instruct_70k_unfilteredThis dataset is the WizardLM dataset victor123/evol_instruct_70k, removing instances of blatant alignment.
54974 instructions remain.
inspired by https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered
All credit to anon8231489123 for the cleanup script that I adapted to wizardlm_clean.py
license: apache-2.0
language:
- en
pretty_name: wizardlm-unfiltered
ru_turbo_alpaca_evol_instructFrench-Alpaca-dataset-Instruct-110K110368 French instructions generated by OpenAI GPT-3.5-turbo in Alpaca Format to finetune general models
Created by Jonathan Pacifico, 2024Please credit my name if you use this dataset in your project.
French-Alpaca-dataset-Instruct-55K55184 french instructions generated by OpenAI GPT-3.5
in Alpaca Format to finetune general models
Created by Jonathan Pacifico
license: apache-2.0
Please credit my name if you use this dataset in your project.
alpaca-instruct-ind-instructionretrieval
alpaca-instruct-ind-instructionretrieval
Deduplicated copy of kornwtp/alpaca-instruct-ind-instructionretrieval,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/alpaca-instruct-ind-instructionretrieval
Deduplicated on: 2026-09-04
Task type: instruction_retrieval
Splits: train
What changed
Kept in this dataset's ORIGINAL schema -- same columns, including the fields the retrieval view discards (Input, output, type, rating). Outputs differing only in… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/alpaca-instruct-ind-instructionretrieval.
