no-refusals
smoltalk-smol-magpie-ultra-no-refusals
SmolTalk Smol-Magpie-Ultra No Refusals
A Minos-cleaned version of HuggingFaceTB/smoltalk / smol-magpie-ultra for use as a neutral helpfulness SFT anchor.
Rows are removed when NousResearch/Minos-v1 classifies the conversation as a refusal. The original train/test split structure is preserved.
Cleaning version: minos-only-v1-2026-06-23
Counts
Split
Input rows
Kept rows
Dropped rows
train
409,537
408,447
1,090
test
21,555
21,488
67
Overall removal… See the full description on the dataset page: https://huggingface.co/datasets/nchapman/smoltalk-smol-magpie-ultra-no-refusals.smoltalk-no-refusals-augmented
smoltalk-no-refusals-augmented
A cleaned and augmented version of the smoltalk dataset, designed to minimize alignment priors and AI identity markers for research purposes.
Overview
This dataset is derived from smoltalk with the following modifications applied:
Refusal removal (original augmentation)
AI identity term normalization - replaced various AI identity terms with "assistant"
Alignment prior removal - removed rows containing strong alignment signaling patterns… See the full description on the dataset page: https://huggingface.co/datasets/EternalRecursion/smoltalk-no-refusals-augmented.Intel-DPO-Pairs-Norefusalsm-a-p_CodeFeedback_norefusals_ShareGPTsmoltalk-no-refusalschat-alpaca-pl-no-refusalsDataset from https://github.com/cascip/ChatAlpaca translated using Hermes 3 405B Instruct(openrouter free), then all the refusals and moralizing was removed. To be used to train chat models in Polish language.
