Arko007/zenyx-v2-SFT-dataset
Zenyx V2 — Raw SFT Dataset Collection This is the unified raw dataset collection used for training Zenyx V2, a custom large language model built from scratch with a novel architecture. Dataset Sources Dataset Rows Category nemotron_sft_code 10,108,883 Code nemotron_sft_math 22,066,397 Math nemotron_sft_science 708,920 Science nemotron_sft_chat 39,792 Chat nemotron_sft_safety 31,426 Safety nemotron_rl 56,339 Instruction Following (RL)… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/zenyx-v2-SFT-dataset.
Zenyx V2 — Raw SFT Dataset Collection
This is the unified raw dataset collection used for training Zenyx V2, a custom large language model built from scratch with a novel architecture.
Dataset Sources
Column Schemas
Nemotron SFT — input, output, category, license, reasoning, generator, used_in_training, version, system_prompt
Nemotron Cascade — domain, source, messages, generator
RedMod Math — text
About Zenyx V2
Zenyx is a custom LLM built with a novel architecture featuring:
- Custom tokenizer
- Modified attention mechanism
- Trained entirely on curated open-source data
This dataset is for research purposes. All source datasets retain their original licenses.
Missing (Next Session)
cascade_chat(~200GB, download interrupted)openO1(corrupt JSON issue, fix pending)stepfun_sft(OOM issue, fix pending)
