datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Open-ert-small-datasetThis is a subset of:
https://huggingface.co/datasets/openerotica/long-roleplay-v0.1
I am using mistral's new DEVSTRAL model to take the entire conversation in JSON format and rate it. I chose DEVSTRAL due to the mistral models being very consistent and well rounded. The Devstral model I was hoping could understand the JSON format a bit better.
I ask the mode to rate each RP based on many different factors including grammar, prose, length (And a few others I will keep to myself :D). I then… See the full description on the dataset page: https://huggingface.co/datasets/SuperbEmphasis/Open-ert-small-dataset.Mini-Mixed-Thoughts
Mini-Mixed-Thoughts
Mini-Mixed-Thoughts is a small, high-density multi-domain reasoning and structured JSON fine-tuning dataset tailored for language models.
This dataset is not made by ertghiu256. It is a combination of open-source available datasets:
Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500
allenai/Dolci-Think-SFT-7B
r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation
open-r1/Mixture-of-Thoughts
Custom synthetic DeepSeek-generated JSON reasoning traces from… See the full description on the dataset page: https://huggingface.co/datasets/ertghiu256/Mini-Mixed-Thoughts.safety-training-distilled-50-examples50 high quality refusal samples distilled from gemini 3.5 flash and deepseek V4 flash.
Prompts are sourced from nvidia/Aegis-AI-Content-Safety-Dataset-2.0 and deepseek v4. Some are crafted manually to mimic natural dangerous questions humans asks.
AtaturkWorldStreamerAtamDatasetreasoning-instruct-distill-mixAtaturkAniDatasetAtamNew
