CoolFace
Datasetpublic

Arko007/zenyx-v3-synthetic-sft-v4

Zenyx-V3-Synthetic-SFT-V4 (Master Expansion) This is the finalized V4 Master Collection for the Zenyx project, expanding upon the previous V3. This version focuses on high-reasoning, code, and math capabilities through massive distillation and thinking-chain integration. Dataset Summary Total Samples: 101,523 Branding: Rebranded to Zenyx / Zenyx Lab. Filtering: Strict English, Math, and Code filter applied. Chinese and non-standard characters removed. Format:… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/zenyx-v3-synthetic-sft-v4.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes1kdownloads
Dataset Card

Zenyx-V3-Synthetic-SFT-V4 (Master Expansion)

This is the finalized V4 Master Collection for the Zenyx project, expanding upon the previous V3. This version focuses on high-reasoning, code, and math capabilities through massive distillation and thinking-chain integration.

Dataset Summary

  • —Total Samples: 101,523
  • —Branding: Rebranded to Zenyx / Zenyx Lab.
  • —Filtering: Strict English, Math, and Code filter applied. Chinese and non-standard characters removed.
  • —Format: Standardized SFT format (Dialogue list with user/assistant roles).

Merged Sources

  1. 1.WithinUsAI/GPT_5.5_Distilled
  2. 2.sequelbox/Titanium4-DeepSeek-V4-Pro
  3. 3.a-m-team/AM-Thinking-v1-RL-Dataset
  4. 4.a-m-team/AM-Thinking-v1-Distilled
  5. 5.a-m-team/AM-DeepSeek-R1-Distilled-1.4M (Curated subset)

Sharding Strategy

The dataset is consolidated into 10 large JSONL shards (train_zenyx_v4_00 to train_zenyx_v4_09) to optimize loading performance for distributed training scripts.