CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01unsloth /alpaca-cleaned Dataset Card for Alpaca-Cleaned Forked from https://huggingface.co/datasets/yahma/alpaca-cleaned Repository: https://github.com/gururise/AlpacaDataCleaned Dataset Description This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset: Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused… See the full description on the dataset page: https://huggingface.co/datasets/unsloth/alpaca-cleaned.texttext-generation10K<n<100K25 likes3.4k downloads9mo agoHugging Face02henry1477 /pcbslm-static-v2-unsloth-vlm PCBSLM static-v2 Unsloth VLM Portable multimodal Unsloth dataset for PCB layout/document-grounded training. The JSONL splits use Unsloth/Gemma-style chat messages: { "messages": [ {"role": "user", "content": [ {"type": "image", "image": "assets/raw_docs/.../images/page.png"}, {"type": "text", "text": "instruction..."} ]}, {"role": "assistant", "content": [ {"type": "text", "text": "{...json answer...}"} ]} ] } Files… See the full description on the dataset page: https://huggingface.co/datasets/henry1477/pcbslm-static-v2-unsloth-vlm.documentimage-text-to-text1K<n<10K0 likes185 downloads5mo agoHugging Face03beezza /ogiri-bokete-unsloth-vlm Japanese Bokete Ogiri — Unsloth VLM format YANS-official/ogiri-bokete を、UnslothのVision SFTで扱える会話形式に変換した非公開用データセットです。 各JSONLレコードは「1画像 + 1回答」です。 { "messages": [ {"role": "user", "content": [ {"type": "image", "image": "images/124469.jpg"}, {"type": "text", "text": "この画像のお題に対して、面白い一言を1つ返してください。"} ]}, {"role": "assistant", "content": [ {"type": "text", "text": "..."} ]} ] } Files train.jsonl: 1,678 records / 630 prompts… See the full description on the dataset page: https://huggingface.co/datasets/beezza/ogiri-bokete-unsloth-vlm.imageimage-to-text1K<n<10K0 likes89 downloads2mo agoHugging Face04r0b0tlab /Hermes-OmniForge-Qwen36-27B-full-v0.3.0-unsloth Hermes OmniForge Qwen3.6-27B Dataset v0.3.0 This package contains the Hermes OmniForge Qwen3.6-27B v0.3.0 synthetic SFT dataset and Unsloth-ready exports. data/final/train.jsonl data/final/validation.jsonl data/final/test.jsonl data/final/*_unsloth_text.jsonl data/final/*_unsloth_vision.jsonl scripts/export_unsloth.py scripts/validate_dataset.py scripts/train_unsloth_text_example.py scripts/train_unsloth_vision_example.py reports/dataset_report.json Dataset Shape… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/Hermes-OmniForge-Qwen36-27B-full-v0.3.0-unsloth.texttext-generation100K<n<1M8 likes83 downloads5mo agoHugging Face05nezahatkorkmaz /unsloth-pmc-vqa-trtext100K<n<1M0 likes77 downloads1y agoHugging Face06malaiwah /glm53-fidelity-gguf-unsloth-udq4kxl-v1 fidelity--glm53.malaiwah.quant.gguf-unsloth-udq4kxl A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from unsloth/GLM-5.3-GGUF. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it).… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm53-fidelity-gguf-unsloth-udq4kxl-v1.tabularn<1K0 likes62 downloads16d agoHugging Face07malaiwah /glm52-fidelity-gguf-unsloth-udq4kxl-v1 fidelity--glm52.malaiwah.quant.gguf-unsloth-udq4kxl A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from unsloth/GLM-5.2-GGUF. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it).… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm52-fidelity-gguf-unsloth-udq4kxl-v1.tabularn<1K0 likes57 downloads16d agoHugging Face08llmguy342 /RealMythosReasoning-unsloth-studio-fixedThe exactly same dataset as RealMythosReasoning(https://huggingface.co/datasets/RealMythos/RealMythosReasoning) but fixed for unsloth studio. Tested on cli and gui, it does work perfectly fine. text1K<n<10K0 likes56 downloads2d agoHugging Face09lakhera2023 /dockerNLcommands-sft-unsloth Docker NL Commands (Unsloth-ready) Converted from dockerNLcommands-sft-jsonl for Unsloth Studio. Configurations Config Format Columns Train rows Test rows alpaca (default) Alpaca instruction, input, output 2294 121 chatml ChatML messages (with role + content) 2294 121 Usage Unsloth Studio Dataset → Hugging Face → lakhera2023/dockerNLcommands-sft-unsloth Format: alpaca (default config) Train split: train, eval split: test… See the full description on the dataset page: https://huggingface.co/datasets/lakhera2023/dockerNLcommands-sft-unsloth.texttext-generation1K<n<10K0 likes50 downloads4mo agoHugging Face10open-llm-leaderboard /sonthenguyen__ft-unsloth-zephyr-sft-bnb-4bit-20241014-170522-detailsgated Dataset Card for Evaluation run of sonthenguyen/ft-unsloth-zephyr-sft-bnb-4bit-20241014-170522 Dataset automatically created during the evaluation run of model sonthenguyen/ft-unsloth-zephyr-sft-bnb-4bit-20241014-170522 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sonthenguyen__ft-unsloth-zephyr-sft-bnb-4bit-20241014-170522-details.tabular10K<n<100K0 likes46 downloads2y agoHugging Face11combatsolutions /tenderdataset_unsloth Dataset Card for Tender Categorization Dataset Dataset Description The Tender Categorization Dataset is designed for categorizing various types of tenders into predefined categories. The dataset includes a diverse set of tender descriptions from different domains, and each entry is categorized according to its type. This dataset is useful for training models to automatically classify tenders based on their descriptions. Data Fields input: The text description… See the full description on the dataset page: https://huggingface.co/datasets/combatsolutions/tenderdataset_unsloth.text100K<n<1M3 likes42 downloads2y agoHugging Face12open-llm-leaderboard /FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face13open-llm-leaderboard /FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face14Skorcht /unslothisstupuidtextn<1K0 likes40 downloads2y agoHugging Face15NeuralNovel /Unsloth-DPO Unsloth-DPO Creator: NeuralNovel Community Organization: ConvexAI Discord: Join us on Discord Special Thanks: Unsloth.ai About Neural-DPO: The Unsloth-DPO dataset, inspired by orca_dpo_pairs. This dataset features questions and answers pairs, with a direct focus on Unsloth.ai. Source Data: orca_dpo_pairs (Inspiration) Make LLM Fine-tuning 2x faster with Unsloth and 🤗… See the full description on the dataset page: https://huggingface.co/datasets/NeuralNovel/Unsloth-DPO.textn<1K6 likes31 downloads3y agoHugging Face16mudler /glaive-unsloth-localaitext10K<n<100K0 likes23 downloads2y agoHugging Face17ThakrePranjal /pharma-preference-dataset-unsloth Pharma DPO Preference Dataset — Unsloth Pipeline Preference dataset in DPO format (prompt / chosen / rejected) used for Stage 3 DPO training in the Unsloth 3-stage pharma fine-tuning pipeline. Format { "prompt": "### Instruction:\nExplain the mechanism of metformin.\n\n### Response:", "chosen": "Metformin primarily acts by activating AMPK...", "rejected": "Metformin mainly works by increasing insulin secretion..." } Stats Total rows: 48… See the full description on the dataset page: https://huggingface.co/datasets/ThakrePranjal/pharma-preference-dataset-unsloth.textn<1K0 likes20 downloads3mo agoHugging Face18mudler /open-o1-sft-unslothtext10K<n<100K0 likes17 downloads2y agoHugging Face19sasindumalhara /unsloth_train_Opus-4.6-Reasoning-24ktext10K<n<100K1 likes17 downloads4mo agoHugging Face20sasindumalhara /unsloth_train_code_train.jsonltext1K<n<10K0 likes16 downloads4mo agoHugging Face21bogdanrivera /legal_civiles_oaxaca_llama_unsloth_template Dataset Card for Código Civil de Oaxaca (Enriched for LLMs) Dataset Description Este es un dataset completo y enriquecido del Código Civil para el Estado de Oaxaca, México, formateado específicamente para el fine-tuning (SFT) de modelos de lenguaje grandes (LLMs) conversacionales. El dataset no solo contiene el texto de los artículos legales, sino que ha sido procesado y aumentado en varias etapas para crear un recurso de alta calidad. El proceso incluye limpieza profunda… See the full description on the dataset page: https://huggingface.co/datasets/bogdanrivera/legal_civiles_oaxaca_llama_unsloth_template.text10K<n<100K0 likes15 downloads11mo agoHugging Face22Benz003 /alpaca-unslothtext10K<n<100K0 likes14 downloads2y agoHugging Face23open-llm-leaderboard /LimYeri__CodeMind-Llama3-8B-unsloth_v4-one-DPO-merged-detailsgated Dataset Card for Evaluation run of LimYeri/CodeMind-Llama3-8B-unsloth_v4-one-DPO-merged Dataset automatically created during the evaluation run of model LimYeri/CodeMind-Llama3-8B-unsloth_v4-one-DPO-merged The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LimYeri__CodeMind-Llama3-8B-unsloth_v4-one-DPO-merged-details.tabular10K<n<100K0 likes14 downloads2y agoHugging Face24Enigmah /Nascenia_Gemma2_Unsloth_Trainingtabularn<1K0 likes13 downloads1mo agoHugging Face25samirun974 /basketeurope_unslothtext10K<n<100K1 likes12 downloads2y agoHugging Face26DerekAUS /Unslothtextn<1K0 likes11 downloads2y agoHugging Face27open-llm-leaderboard /FlofloB__10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face28open-llm-leaderboard /unsloth__Phi-3-mini-4k-instruct-detailsgated Dataset Card for Evaluation run of unsloth/Phi-3-mini-4k-instruct Dataset automatically created during the evaluation run of model unsloth/Phi-3-mini-4k-instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/unsloth__Phi-3-mini-4k-instruct-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face29open-llm-leaderboard /unsloth__phi-4-bnb-4bit-detailsgated Dataset Card for Evaluation run of unsloth/phi-4-bnb-4bit Dataset automatically created during the evaluation run of model unsloth/phi-4-bnb-4bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/unsloth__phi-4-bnb-4bit-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face30open-llm-leaderboard /unsloth__phi-4-unsloth-bnb-4bit-detailsgated Dataset Card for Evaluation run of unsloth/phi-4-unsloth-bnb-4bit Dataset automatically created during the evaluation run of model unsloth/phi-4-unsloth-bnb-4bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/unsloth__phi-4-unsloth-bnb-4bit-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.