usamakenway/unified-uncensored-qwen-chatml-sft
Unified Uncensored Qwen SFT Dataset This dataset is a mixed-license compilation of instruction/chat datasets converted into a single Qwen/ChatML-style text JSONL format. Format Each row has: { "text": "<|im_start|>user\n...<|im_end|>\n<|im_start|>assistant\n...<|im_end|>", "source": "dataset/repo", "source_format": "alpaca|sharegpt|messages|human_bot_text|prompt_response", "source_license": "apache-2.0|mit|cc-by-4.0|cc-by-nc-4.0|other|unknown"… See the full description on the dataset page: https://huggingface.co/datasets/usamakenway/unified-uncensored-qwen-chatml-sft.
Unified Uncensored Qwen SFT Dataset
This dataset is a mixed-license compilation of instruction/chat datasets converted into a single Qwen/ChatML-style text JSONL format.
Format
Each row has:
{
"text": "<|im_start|>user\n...<|im_end|>\n<|im_start|>assistant\n...<|im_end|>",
"source": "dataset/repo",
"source_format": "alpaca|sharegpt|messages|human_bot_text|prompt_response",
"source_license": "apache-2.0|mit|cc-by-4.0|cc-by-nc-4.0|other|unknown",
"source_family": "..."
}License
This is a mixed-license compilation. Each row preserves original source and license metadata. Users are responsible for complying with the terms of each source dataset.
Do not assume the full combined dataset is commercially usable. For commercial use, create a filtered export that only includes sources with explicit permissive licenses and review attribution requirements.
Processing
Rows were normalized into intermediate messages, rendered into Qwen/ChatML text, and deduplicated with exact and near-duplicate checks.
See the generated sources_manifest.json and dedup_report.json for source-level counts, licenses, and conversion statistics.
