CoolFace
Datasetpublic

SMH-DEV-AI/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54

NOESIS DORA SFT Dataset Multilingual supervised fine-tuning dataset built for the NOESIS QwQ+DeepSeek-R1 MoE pipeline. Released as part of the NOESIS Professional Multilingual Dubbing Automation Platform(framework: DHCF-FNO — Deterministic Hybrid Control Framework for Frozen Neural Operators). Founder: Ilia Bolotnikov Organization: AMAImedia.com X (Twitter): @AMAImediacom LinkedIn: Ilia Bolotnikov Telegram: @djbionicl NOESIS version: v14.8-NT89 Build date: 2026-04… See the full description on the dataset page: https://huggingface.co/datasets/SMH-DEV-AI/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
1likes56downloads
Dataset Card

NOESIS DORA SFT Dataset

Multilingual supervised fine-tuning dataset built for the NOESIS QwQ+DeepSeek-R1 MoE pipeline.

Released as part of the NOESIS Professional Multilingual Dubbing Automation Platform (framework: DHCF-FNO — Deterministic Hybrid Control Framework for Frozen Neural Operators).


Files

FileRecordsSizeDescription
NOESIS-1M-multilingual-reasoning-router-general-code-math-psych-aya-sft-claude-sonet46-opus47-deepseek4-qwen36-gemini31-r1-gpt54.jsonl1,000,000~1.4 GBFull 1M SFT dataset (2026-04-18)
NOESIS-50K-multilingual-reasoning-router-general-code-math-psych-aya-sft-claude-sonet46-opus47-deepseek4-qwen36-gemini31-r1-gpt54.jsonl50,000~66 MB50K curated high-quality subset, top-30 language quotas

Format

All files are JSONL — one JSON object per line:

json
{"text": "User: <question>\nAssistant: <answer>"}

Records with <think>...</think> blocks contain reasoning traces from QwQ-32B / DeepSeek-R1 heritage.


Dataset composition (1M)

SourceRecordsNotes
Aya dataset (Cohere, 204k)~197,000Multilingual instruction, 101 languages
Claude Sonnet 4.6 SFT~122,000High-quality EN assistant turns
DeepSeek-R1-Distill-7B synthetic~41,000Reasoning traces with <think>
NOESIS translation pairs (50k)~46,00030-language parallel SFT
Claude Opus 4.7 thinking~25,000Extended reasoning traces
Other SFT sources~36,000Code, math, research
Additional mixed sources~533,000Rebalanced multilingual SFT

50k curated selection strategy

The NOESIS-50K-multilingual-...-gpt54.jsonl is sampled from the 1M with quality scoring:

Quality score (higher = selected first):

  • —+3 if <think>...</think> present (reasoning trace)
  • —+2 if assistant response > 2000 chars
  • —+1 if assistant response > 500 chars
  • —+1 if contains code (``` or def/function/class)
  • —+1 if contains math (LaTeX symbols, ∑ ∫ ≤ ≥)

Language quotas (35,000 total across top-30 languages):

LangQuotaLangQuotaLangQuota
EN8,000ID1,000UK500
ZH4,000DE1,000PL500
HI2,500JA1,000NL500
ES2,500KO800TA500
AR2,000TR800MS400
FR2,000VI800SW400
RU1,500FA700HA400
PT1,500IT700GU400
BN600KK400
TH600UZ400
MR400
UR400

English high-quality pool: 15,000 records (reasoning/code/math priority)


Contributing AI models

Synthetic SFT records in this dataset were generated by or distilled from outputs of:

ModelUsage
Claude Sonnet 4.6High-quality EN instruction, coding, analysis
Claude Opus 4.7 (thinking)Extended reasoning traces
DeepSeek-R1 / R1-Distill<think> reasoning chain records
DeepSeek V4General instruction, coding, and reasoning
Qwen3.6Multilingual and reasoning SFT
Gemini 3.1General instruction and research
GPT-5.4Diverse instruction-following turns

Intended use

This dataset is designed for:

  • —DoRA SFT fine-tuning of Qwen3-based MoE models
  • —Router fine-tuning (gate.weight training) for CMoE architectures
  • —Multilingual instruction tuning with reasoning trace distillation

Primary target: NOESIS-QwQ-R1 pipeline (QwQ-32B + DeepSeek-R1-32B TIES merge → CMoE 16E).


License

Apache License 2.0.

Dataset composition includes records derived from:

  • —Aya dataset — Apache 2.0, Cohere
  • —Original NOESIS synthetic data — Apache 2.0, AMAImedia.com 2026

See LICENSE file for full terms.


HuggingFace repos

DatasetHuggingFace repo
1M fullAMAImedia/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54
50K curatedAMAImedia/NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54

Note: HuggingFace enforces a 96-character repo ID limit. The full dataset name is encoded in the filename.


Citation

bibtex
@misc{noesis_dora_dataset_2026,
  title  = {NOESIS DORA SFT Dataset — 1M multilingual instruction-tuning records},
  author = {Bolotnikov, Ilia},
  year   = {2026},
  publisher = {AMAImedia},
  url    = {https://amaimedia.com}
}