unsloth
Datasets
All datasets matching “unsloth”alpaca-cleaned
Dataset Card for Alpaca-Cleaned
Forked from https://huggingface.co/datasets/yahma/alpaca-cleaned
Repository: https://github.com/gururise/AlpacaDataCleaned
Dataset Description
This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused… See the full description on the dataset page: https://huggingface.co/datasets/unsloth/alpaca-cleaned.LaTeX_OCR1% sampled from https://huggingface.co/datasets/linxy/LaTeX_OCR
OpenMathReasoning-miniRadiology_mini0.33% sampled from https://huggingface.co/datasets/eltorio/ROCOv2-radiology
llava-instruct-mix-vsft-miniOriginally from https://huggingface.co/datasets/HuggingFaceH4/llava-instruct-mix-vsft but 0.33% randomnly sampled
receipted-unsloth
Receipted Unsloth
How SZL Holdings actually trains. Silhouette from Unsloth QLoRA. Cut is original SZL. We do not republish Unsloth Studio, Desktop, copy, code, or someone else's tensors.
Collection: Receipted Unsloth — LIVE
The house loop
Disclose the Apache base (Qwen/Qwen2.5-* or Qwen/Qwen3.5-0.8B).
Train with Unsloth FastLanguageModel QLoRA on owner metal or HF Jobs (uv run + HF_TOKEN).
Bind dataset SHA-256, LoRA knobs, seed, and loss into a training receipt.… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/receipted-unsloth.
