Podtech/Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt
Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt Dataset Overview This dataset is a reformatted subset of the tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1 dataset, specifically derived from the v1-Ja-202601 subset. It was created to facilitate Continuous Pre-Training (CPT) by extracting only the text_gpt_oss field from the original data. Dataset Statistics & Token Counts The token counts for each category were calculated using the… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt.
Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt
Dataset Overview
This dataset is a reformatted subset of the tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1 dataset, specifically derived from the v1-Ja-202601 subset. It was created to facilitate Continuous Pre-Training (CPT) by extracting only the text_gpt_oss field from the original data.
Dataset Statistics & Token Counts
The token counts for each category were calculated using the gigatoken library with the openai/gpt-oss-20b tokenizer.
License
This dataset is licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0) available at https://creativecommons.org/licenses/by/4.0/legalcode.
