CoolFace
Datasetpublic

Podtech/Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt

Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt Dataset Overview This dataset is a reformatted subset of the tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1 dataset, specifically derived from the v1-Ja-202601 subset. It was created to facilitate Continuous Pre-Training (CPT) by extracting only the text_gpt_oss field from the original data. Dataset Statistics & Token Counts The token counts for each category were calculated using the… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt.

sourceHugging Facecc-by-4.0updated 1mo agoView on Hugging Face
0likes444downloads
Dataset Card

Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt

Dataset Overview

This dataset is a reformatted subset of the tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1 dataset, specifically derived from the v1-Ja-202601 subset. It was created to facilitate Continuous Pre-Training (CPT) by extracting only the text_gpt_oss field from the original data.

Dataset Statistics & Token Counts

The token counts for each category were calculated using the gigatoken library with the openai/gpt-oss-20b tokenizer.

CategoryFile CountTotal TokensToken Size
Code18 files6,285,974,424~6.29 B
Math19 files6,927,925,779~6.93 B
STEM24 files3,205,527,028~3.21 B

License

This dataset is licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0) available at https://creativecommons.org/licenses/by/4.0/legalcode.