SMH-DEV-AI/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54
NOESIS DORA SFT Dataset Multilingual supervised fine-tuning dataset built for the NOESIS QwQ+DeepSeek-R1 MoE pipeline. Released as part of the NOESIS Professional Multilingual Dubbing Automation Platform(framework: DHCF-FNO — Deterministic Hybrid Control Framework for Frozen Neural Operators). Founder: Ilia Bolotnikov Organization: AMAImedia.com X (Twitter): @AMAImediacom LinkedIn: Ilia Bolotnikov Telegram: @djbionicl NOESIS version: v14.8-NT89 Build date: 2026-04… See the full description on the dataset page: https://huggingface.co/datasets/SMH-DEV-AI/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.
NOESIS DORA SFT Dataset
Multilingual supervised fine-tuning dataset built for the NOESIS QwQ+DeepSeek-R1 MoE pipeline.
Released as part of the NOESIS Professional Multilingual Dubbing Automation Platform (framework: DHCF-FNO — Deterministic Hybrid Control Framework for Frozen Neural Operators).
- Founder: Ilia Bolotnikov
- Organization: AMAImedia.com
- X (Twitter): @AMAImediacom
- LinkedIn: Ilia Bolotnikov
- Telegram: @djbionicl
- NOESIS version: v14.8-NT89
- Build date: 2026-04
Files
Format
All files are JSONL — one JSON object per line:
{"text": "User: <question>\nAssistant: <answer>"}Records with <think>...</think> blocks contain reasoning traces from QwQ-32B / DeepSeek-R1 heritage.
Dataset composition (1M)
50k curated selection strategy
The NOESIS-50K-multilingual-...-gpt54.jsonl is sampled from the 1M with quality scoring:
Quality score (higher = selected first):
- +3 if
<think>...</think>present (reasoning trace) - +2 if assistant response > 2000 chars
- +1 if assistant response > 500 chars
- +1 if contains code (``` or def/function/class)
- +1 if contains math (LaTeX symbols, ∑ ∫ ≤ ≥)
Language quotas (35,000 total across top-30 languages):
English high-quality pool: 15,000 records (reasoning/code/math priority)
Contributing AI models
Synthetic SFT records in this dataset were generated by or distilled from outputs of:
Intended use
This dataset is designed for:
- DoRA SFT fine-tuning of Qwen3-based MoE models
- Router fine-tuning (gate.weight training) for CMoE architectures
- Multilingual instruction tuning with reasoning trace distillation
Primary target: NOESIS-QwQ-R1 pipeline (QwQ-32B + DeepSeek-R1-32B TIES merge → CMoE 16E).
License
Apache License 2.0.
Dataset composition includes records derived from:
- Aya dataset — Apache 2.0, Cohere
- Original NOESIS synthetic data — Apache 2.0, AMAImedia.com 2026
See LICENSE file for full terms.
HuggingFace repos
Note: HuggingFace enforces a 96-character repo ID limit. The full dataset name is encoded in the filename.
Citation
@misc{noesis_dora_dataset_2026,
title = {NOESIS DORA SFT Dataset — 1M multilingual instruction-tuning records},
author = {Bolotnikov, Ilia},
year = {2026},
publisher = {AMAImedia},
url = {https://amaimedia.com}
}