Arko007/zenyx-v3-synthetic-sft-v4
Zenyx-V3-Synthetic-SFT-V4 (Master Expansion) This is the finalized V4 Master Collection for the Zenyx project, expanding upon the previous V3. This version focuses on high-reasoning, code, and math capabilities through massive distillation and thinking-chain integration. Dataset Summary Total Samples: 101,523 Branding: Rebranded to Zenyx / Zenyx Lab. Filtering: Strict English, Math, and Code filter applied. Chinese and non-standard characters removed. Format:… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/zenyx-v3-synthetic-sft-v4.
Zenyx-V3-Synthetic-SFT-V4 (Master Expansion)
This is the finalized V4 Master Collection for the Zenyx project, expanding upon the previous V3. This version focuses on high-reasoning, code, and math capabilities through massive distillation and thinking-chain integration.
Dataset Summary
- Total Samples: 101,523
- Branding: Rebranded to Zenyx / Zenyx Lab.
- Filtering: Strict English, Math, and Code filter applied. Chinese and non-standard characters removed.
- Format: Standardized SFT format (Dialogue list with user/assistant roles).
Merged Sources
- WithinUsAI/GPT_5.5_Distilled
- sequelbox/Titanium4-DeepSeek-V4-Pro
- a-m-team/AM-Thinking-v1-RL-Dataset
- a-m-team/AM-Thinking-v1-Distilled
- a-m-team/AM-DeepSeek-R1-Distilled-1.4M (Curated subset)
Sharding Strategy
The dataset is consolidated into 10 large JSONL shards (train_zenyx_v4_00 to train_zenyx_v4_09) to optimize loading performance for distributed training scripts.
