datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rosetta-ko-math-synth-sft
rosetta-ko-math-synth-sft
Korean-native mathematics data — problems with step-by-step Korean solutions, answer-verified by Math-Verify (LaTeX \boxed{} + symbolic equivalence).
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Math Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-math-synth-sft (this repo)
supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-math-synth-sft.rosetta-ko-chat-synth-sft
rosetta-ko-chat-synth-sft
Korean-native general-purpose conversational data — diverse everyday instructions with natural Korean answers.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Chat Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-chat-synth-sft (this repo)
supervised fine-tuning (user/assistant messages)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-chat-synth-sft.rosetta-ko-law-synth-sft
rosetta-ko-law-synth-sft
Korean legal-domain data grounded in national statutes — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Law Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-law-synth-cpt
continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-law-synth-sft.rosetta-ko-code-synth-sft
rosetta-ko-code-synth-sft
Korean-native coding data — problems with solutions whose unit tests were actually executed and passed (execution-grounded rejection sampling).
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Code Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-code-synth-sft (this repo)
supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-code-synth-sft.rosetta-ko-tourism-synth-sft
rosetta-ko-tourism-synth-sft
Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Tourism Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-tourism-synth-cpt
continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-sft.rosetta-ko-heritage-synth-sft
rosetta-ko-heritage-synth-sft
Korean cultural-heritage data grounded in public heritage records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Cultural Heritage Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-heritage-synth-cpt
continued-pretraining corpus… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-heritage-synth-sft.rosetta-ko-instruction-following-synth-sft
rosetta-ko-instruction-following-synth-sft
Korean-native instruction-following data — instructions with verifiable constraints (length, format, keywords, JSON, ...) checked by programmatic verifiers.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Instruction-Following Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-instruction-following-synth-sft.
