datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alpaca-cleaned
Dataset Card for Alpaca-Cleaned
Forked from https://huggingface.co/datasets/yahma/alpaca-cleaned
Repository: https://github.com/gururise/AlpacaDataCleaned
Dataset Description
This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused… See the full description on the dataset page: https://huggingface.co/datasets/unsloth/alpaca-cleaned.pcbslm-static-v2-unsloth-vlm
PCBSLM static-v2 Unsloth VLM
Portable multimodal Unsloth dataset for PCB layout/document-grounded training.
The JSONL splits use Unsloth/Gemma-style chat messages:
{
"messages": [
{"role": "user", "content": [
{"type": "image", "image": "assets/raw_docs/.../images/page.png"},
{"type": "text", "text": "instruction..."}
]},
{"role": "assistant", "content": [
{"type": "text", "text": "{...json answer...}"}
]}
]
}
Files… See the full description on the dataset page: https://huggingface.co/datasets/henry1477/pcbslm-static-v2-unsloth-vlm.ogiri-bokete-unsloth-vlm
Japanese Bokete Ogiri — Unsloth VLM format
YANS-official/ogiri-bokete を、UnslothのVision SFTで扱える会話形式に変換した非公開用データセットです。
各JSONLレコードは「1画像 + 1回答」です。
{
"messages": [
{"role": "user", "content": [
{"type": "image", "image": "images/124469.jpg"},
{"type": "text", "text": "この画像のお題に対して、面白い一言を1つ返してください。"}
]},
{"role": "assistant", "content": [
{"type": "text", "text": "..."}
]}
]
}
Files
train.jsonl: 1,678 records / 630 prompts… See the full description on the dataset page: https://huggingface.co/datasets/beezza/ogiri-bokete-unsloth-vlm.Hermes-OmniForge-Qwen36-27B-full-v0.3.0-unsloth
Hermes OmniForge Qwen3.6-27B Dataset v0.3.0
This package contains the Hermes OmniForge Qwen3.6-27B v0.3.0 synthetic SFT dataset and Unsloth-ready exports.
data/final/train.jsonl
data/final/validation.jsonl
data/final/test.jsonl
data/final/*_unsloth_text.jsonl
data/final/*_unsloth_vision.jsonl
scripts/export_unsloth.py
scripts/validate_dataset.py
scripts/train_unsloth_text_example.py
scripts/train_unsloth_vision_example.py
reports/dataset_report.json
Dataset Shape… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/Hermes-OmniForge-Qwen36-27B-full-v0.3.0-unsloth.unsloth-pmc-vqa-trglm53-fidelity-gguf-unsloth-udq4kxl-v1
fidelity--glm53.malaiwah.quant.gguf-unsloth-udq4kxl
A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from unsloth/GLM-5.3-GGUF.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it).… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm53-fidelity-gguf-unsloth-udq4kxl-v1.glm52-fidelity-gguf-unsloth-udq4kxl-v1
fidelity--glm52.malaiwah.quant.gguf-unsloth-udq4kxl
A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from unsloth/GLM-5.2-GGUF.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it).… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm52-fidelity-gguf-unsloth-udq4kxl-v1.RealMythosReasoning-unsloth-studio-fixedThe exactly same dataset as RealMythosReasoning(https://huggingface.co/datasets/RealMythos/RealMythosReasoning) but fixed for unsloth studio. Tested on cli and gui, it does work perfectly fine.
dockerNLcommands-sft-unsloth
Docker NL Commands (Unsloth-ready)
Converted from dockerNLcommands-sft-jsonl for Unsloth Studio.
Configurations
Config
Format
Columns
Train rows
Test rows
alpaca (default)
Alpaca
instruction, input, output
2294
121
chatml
ChatML
messages (with role + content)
2294
121
Usage
Unsloth Studio
Dataset → Hugging Face → lakhera2023/dockerNLcommands-sft-unsloth
Format: alpaca (default config)
Train split: train, eval split: test… See the full description on the dataset page: https://huggingface.co/datasets/lakhera2023/dockerNLcommands-sft-unsloth.sonthenguyen__ft-unsloth-zephyr-sft-bnb-4bit-20241014-170522-details
Dataset Card for Evaluation run of sonthenguyen/ft-unsloth-zephyr-sft-bnb-4bit-20241014-170522
Dataset automatically created during the evaluation run of model sonthenguyen/ft-unsloth-zephyr-sft-bnb-4bit-20241014-170522
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sonthenguyen__ft-unsloth-zephyr-sft-bnb-4bit-20241014-170522-details.tenderdataset_unsloth
Dataset Card for Tender Categorization Dataset
Dataset Description
The Tender Categorization Dataset is designed for categorizing various types of tenders into predefined categories. The dataset includes a diverse set of tender descriptions from different domains, and each entry is categorized according to its type. This dataset is useful for training models to automatically classify tenders based on their descriptions.
Data Fields
input: The text description… See the full description on the dataset page: https://huggingface.co/datasets/combatsolutions/tenderdataset_unsloth.FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.unslothisstupuidUnsloth-DPO
Unsloth-DPO
Creator: NeuralNovel
Community Organization: ConvexAI
Discord: Join us on Discord
Special Thanks: Unsloth.ai
About Neural-DPO: The Unsloth-DPO dataset, inspired by orca_dpo_pairs. This dataset features questions and answers pairs, with a direct focus on Unsloth.ai.
Source Data:
orca_dpo_pairs (Inspiration)
Make LLM Fine-tuning 2x faster with Unsloth and 🤗… See the full description on the dataset page: https://huggingface.co/datasets/NeuralNovel/Unsloth-DPO.glaive-unsloth-localaipharma-preference-dataset-unsloth
Pharma DPO Preference Dataset — Unsloth Pipeline
Preference dataset in DPO format (prompt / chosen / rejected) used for
Stage 3 DPO training in the Unsloth 3-stage pharma fine-tuning pipeline.
Format
{
"prompt": "### Instruction:\nExplain the mechanism of metformin.\n\n### Response:",
"chosen": "Metformin primarily acts by activating AMPK...",
"rejected": "Metformin mainly works by increasing insulin secretion..."
}
Stats
Total rows: 48… See the full description on the dataset page: https://huggingface.co/datasets/ThakrePranjal/pharma-preference-dataset-unsloth.open-o1-sft-unslothunsloth_train_Opus-4.6-Reasoning-24kunsloth_train_code_train.jsonllegal_civiles_oaxaca_llama_unsloth_template
Dataset Card for Código Civil de Oaxaca (Enriched for LLMs)
Dataset Description
Este es un dataset completo y enriquecido del Código Civil para el Estado de Oaxaca, México, formateado específicamente para el fine-tuning (SFT) de modelos de lenguaje grandes (LLMs) conversacionales.
El dataset no solo contiene el texto de los artículos legales, sino que ha sido procesado y aumentado en varias etapas para crear un recurso de alta calidad. El proceso incluye limpieza profunda… See the full description on the dataset page: https://huggingface.co/datasets/bogdanrivera/legal_civiles_oaxaca_llama_unsloth_template.alpaca-unslothLimYeri__CodeMind-Llama3-8B-unsloth_v4-one-DPO-merged-details
Dataset Card for Evaluation run of LimYeri/CodeMind-Llama3-8B-unsloth_v4-one-DPO-merged
Dataset automatically created during the evaluation run of model LimYeri/CodeMind-Llama3-8B-unsloth_v4-one-DPO-merged
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LimYeri__CodeMind-Llama3-8B-unsloth_v4-one-DPO-merged-details.Nascenia_Gemma2_Unsloth_Trainingbasketeurope_unslothUnslothFlofloB__10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit-details.unsloth__Phi-3-mini-4k-instruct-details
Dataset Card for Evaluation run of unsloth/Phi-3-mini-4k-instruct
Dataset automatically created during the evaluation run of model unsloth/Phi-3-mini-4k-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/unsloth__Phi-3-mini-4k-instruct-details.unsloth__phi-4-bnb-4bit-details
Dataset Card for Evaluation run of unsloth/phi-4-bnb-4bit
Dataset automatically created during the evaluation run of model unsloth/phi-4-bnb-4bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/unsloth__phi-4-bnb-4bit-details.unsloth__phi-4-unsloth-bnb-4bit-details
Dataset Card for Evaluation run of unsloth/phi-4-unsloth-bnb-4bit
Dataset automatically created during the evaluation run of model unsloth/phi-4-unsloth-bnb-4bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/unsloth__phi-4-unsloth-bnb-4bit-details.
