CoolFace
Datasetpublic

Podtech/Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt

Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt Dataset Overview This dataset is a reformatted subset of the tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1 dataset, specifically derived from the v1-Ja-202601 subset. It was created to facilitate Continuous Pre-Training (CPT) by extracting only the text_gpt_oss field from the original data. Dataset Statistics & Token Counts The token counts for each category were calculated using the… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt.

sourceHugging Facecc-by-4.0updated 1mo agoView on Hugging Face
0likes509downloads
5 commits on main
23883851mo ago

Filter dataset to keep only Japanese queries and answers

KawamuraH
6524fff1mo ago

Delete data

KawamuraH
824d1c22mo ago

Update README

KawamuraH
98a8aef2mo ago

first commit

KawamuraH
14ce7792mo ago

initial commit

KawamuraH