datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
combined-instruct-sft
Combined Instruct SFT Mixture
A curated, globally shuffled mixture of 2,980,737 high-quality instruction-following dialogues harmonized into the standardized OpenAI / ChatML format.
Dataset Mixture Breakdown
Source Dataset
Split & Filtering Restrictions
Harmonized Count
allenai/Dolci-Instruct-SFT
Train split, Safety category samples removed
2,041,725
nvidia/Nemotron-SFT-Instruction-Following-Chat-v2
reasoning_off split, English dialogues only
888,460… See the full description on the dataset page: https://huggingface.co/datasets/DrRiceIO7/combined-instruct-sft.SlopReviewHowdy! This is a curated dataset for training models to distinguish between Slop and Quality Writing. You could also use it to train an LLM to write, but most of the AI responses are cosnidered Slop, so it's not recommended. Might also want to check in with specific models' licenses to see if distillation is allowed.
This dataset was made by feeding 200 prompts from ChaoticNeutrals/Reddit-SFW-Writing_Prompts_ShareGPT into various LLMs.
In v1, I compared the responses with the human generated… See the full description on the dataset page: https://huggingface.co/datasets/DrRiceIO7/SlopReview.AlteredDatasetA dataset combining various different other datasets, including my datasets megamergecreative(reverse prompted from chunked stories, novels, and other data), testingSCP(the amount of data is so small, it probably does nothing, but I hope it adds a sorta background note), and thebigdataset-cleaned(which already contains the other 2 but I'm hoping the extra data will improve the creative writing). Alongside this is data from agentlans/HumanLLMs-Human-Like-DPO-Dataset-no-emojis and… See the full description on the dataset page: https://huggingface.co/datasets/DrRiceIO7/AlteredDataset.sketchy-dataset
Sketchy Dataset
This is the Sketchy database used for sketch-to-image generation research.
Dataset Details
Uploaded: 2026-03-09
Files: 87987
Total Size: 1.03 GB
Format: Images (sketches and photos)
Dataset Structure
sketchy/
├── train/
│ ├── sketches/
│ └── photos/
├── val/
│ ├── sketches/
│ └── photos/
└── test/
├── sketches/
└── photos/
Usage
from datasets import load_dataset
# Load the dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/DrRORAL/sketchy-dataset.aggressive-instructionssplit_datasetCoT-2k-Tokensafrica-unsdg-score-of-adoption-and-implementation-of-national-drr-st-sg-dsr-lgrgsr
Africa Unsdg Score of Adoption and Implementation of National Drr St Sg Dsr Lgrgsr | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: parquet - Sector: other_unclassified - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-unsdg-score-of-adoption-and-implementation-of-national-drr-st-sg-dsr-lgrgsr.thebigdatasetMedical_Customer_carecreativeharmCombined my testingharm and megamergecreative datasets. Surprisingly good result, as seen in mergedhereticFT. Still not perfect though.
dataset_info:
features:
- name: prompt
dtype: string
- name: response
dtype: string
- name: source_id
dtype: string
splits:
- name: train
num_bytes: 56502516.0
num_examples: 84006
download_size: 31997651
dataset_size: 56502516.0
configs:
- config_name: default
data_files:
- split: train
path:… See the full description on the dataset page: https://huggingface.co/datasets/DrRiceIO7/creativeharm.thebigdataset-cleanedCleaned up version of my dataset thebigdataset with empty rows removed.
aggressive_dpo_cleanedaggressive_dpo_reformattedmeanDPOThis dataset is based on anthropic/model-written-evals (specifically the sycophany evals), with a simple python script designed to randomly piece together different phrases to make it either sound mean or really pathetic.
megamergecreativeMy first dataset. It's reverse prompted. It's also bad. Don't use it. You will drive your LLM insane if you try.
testingharmAll the purely harmful responses from PKU-Alignment/PKU-SafeRLHF. Surprisingly good by itself to reduce refusals.
Short-CoTmedibot_dataset_A2nd_datasettestingSCPEFC-miniEgyptain Forums Corpus-mini: A subset of the original EFC corpus used in pretraining EgyBERT. The total size of the mini version is 2.0 GB separated into multiple txt file. Can be downloaded directly from "Files and versions"
asia-unsdg-score-of-adoption-and-implementation-of-national-drr-st-sg-dsr-lgrgsr
