datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rosettacode-parsed
Data Origins
Original dataset: https://huggingface.co/datasets/jondurbin/rosettacode-raw/
Cleaner code: https://github.com/the-crypt-keeper/rosettacode-parser
Data Fields
Field
Type
Description
title
string
problem title
task
string
problem description
language
string
solution language/variant
soulution
string
solution source code
Languages
One .jsonl is provided per language group, the sublanguage field in the data denotes the… See the full description on the dataset page: https://huggingface.co/datasets/mike-ravkine/rosettacode-parsed.rosettabench-150-stratified-compressed
About Dataset
Why?
Used for RosettaBench (A contamination-free benchmark for measuring learning, not memorization).
Preparation
Source: LiveCodeBench (release_v5, AtCoder platform only), accessed via sam-paech/livecodebench-code_generation_lite on Hugging Face.
AtCoder problems were selected for platform consistency and STDIN/STDOUT format compatibility. Problems with fewer than 3 test cases were excluded, leaving a pool of 342 problems.
Sampling: 150… See the full description on the dataset page: https://huggingface.co/datasets/namanbnsl/rosettabench-150-stratified-compressed.rosetta-ko-math-synth-sft
rosetta-ko-math-synth-sft
Korean-native mathematics data — problems with step-by-step Korean solutions, answer-verified by Math-Verify (LaTeX \boxed{} + symbolic equivalence).
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Math Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-math-synth-sft (this repo)
supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-math-synth-sft.rosetta-ko-math-synth-sft-think
rosetta-ko-math-synth-sft-think
Korean-native mathematics data — problems with step-by-step Korean solutions, answer-verified by Math-Verify (LaTeX \boxed{} + symbolic equivalence).
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Math Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-math-synth-sft
supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-math-synth-sft-think.rosetta-ko-chat-synth-sft-think
rosetta-ko-chat-synth-sft-think
Korean-native general-purpose conversational data — diverse everyday instructions with natural Korean answers.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Chat Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-chat-synth-sft
supervised fine-tuning (user/assistant messages)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-chat-synth-sft-think.rosetta-ko-chat-synth-sft
rosetta-ko-chat-synth-sft
Korean-native general-purpose conversational data — diverse everyday instructions with natural Korean answers.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Chat Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-chat-synth-sft (this repo)
supervised fine-tuning (user/assistant messages)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-chat-synth-sft.rosetta-ko-chat-synth-rlvr
rosetta-ko-chat-synth-rlvr
Korean-native general-purpose conversational data — diverse everyday instructions with natural Korean answers.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Chat Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-chat-synth-sft
supervised fine-tuning (user/assistant messages)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-chat-synth-rlvr.rosetta-ko-heritage-synth-cpt
rosetta-ko-heritage-synth-cpt
Korean cultural-heritage data grounded in public heritage records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Cultural Heritage Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-heritage-synth-cpt (this repo)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-heritage-synth-cpt.rosetta-ko-law-synth-sft
rosetta-ko-law-synth-sft
Korean legal-domain data grounded in national statutes — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Law Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-law-synth-cpt
continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-law-synth-sft.rosetta-ko-math-synth-dpo
rosetta-ko-math-synth-dpo
Korean-native mathematics data — problems with step-by-step Korean solutions, answer-verified by Math-Verify (LaTeX \boxed{} + symbolic equivalence).
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Math Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-math-synth-sft
supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-math-synth-dpo.rosetta-ko-math-synth-rlvr
rosetta-ko-math-synth-rlvr
Korean-native mathematics data — problems with step-by-step Korean solutions, answer-verified by Math-Verify (LaTeX \boxed{} + symbolic equivalence).
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Math Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-math-synth-sft
supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-math-synth-rlvr.rosetta-ko-instruction-following-synth-rlvr
rosetta-ko-instruction-following-synth-rlvr
Korean-native instruction-following data — instructions with verifiable constraints (length, format, keywords, JSON, ...) checked by programmatic verifiers.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Instruction-Following Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-instruction-following-synth-rlvr.rosetta-ko-instruction-following-synth-sft-think
rosetta-ko-instruction-following-synth-sft-think
Korean-native instruction-following data — instructions with verifiable constraints (length, format, keywords, JSON, ...) checked by programmatic verifiers.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Instruction-Following Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-instruction-following-synth-sft-think.rosetta-ko-law-synth-cpt
rosetta-ko-law-synth-cpt
Korean legal-domain data grounded in national statutes — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Law Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-law-synth-cpt (this repo)
continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-law-synth-cpt.rosetta-ko-law-synth-rlvr
rosetta-ko-law-synth-rlvr
Korean legal-domain data grounded in national statutes — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Law Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-law-synth-cpt
continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-law-synth-rlvr.rosetta-ko-tourism-synth-cpt
rosetta-ko-tourism-synth-cpt
Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Tourism Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-tourism-synth-cpt (this repo)
continued-pretraining corpus (plain… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-cpt.rosetta-ko-tourism-synth-rlvr
rosetta-ko-tourism-synth-rlvr
Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Tourism Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-tourism-synth-cpt
continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-rlvr.rosetta-ko-code-synth-sft
rosetta-ko-code-synth-sft
Korean-native coding data — problems with solutions whose unit tests were actually executed and passed (execution-grounded rejection sampling).
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Code Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-code-synth-sft (this repo)
supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-code-synth-sft.rosetta-ko-heritage-synth-rlvr
rosetta-ko-heritage-synth-rlvr
Korean cultural-heritage data grounded in public heritage records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Cultural Heritage Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-heritage-synth-cpt
continued-pretraining corpus… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-heritage-synth-rlvr.rosetta-ko-heritage-synth-sft-think
rosetta-ko-heritage-synth-sft-think
Korean cultural-heritage data grounded in public heritage records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Cultural Heritage Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-heritage-synth-cpt
continued-pretraining… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-heritage-synth-sft-think.rosetta-ko-law-synth-sft-think
rosetta-ko-law-synth-sft-think
Korean legal-domain data grounded in national statutes — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Law Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-law-synth-cpt
continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-law-synth-sft-think.rosetta-ko-math-synth-dpo-think
rosetta-ko-math-synth-dpo-think
Korean-native mathematics data — problems with step-by-step Korean solutions, answer-verified by Math-Verify (LaTeX \boxed{} + symbolic equivalence).
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Math Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-math-synth-sft
supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-math-synth-dpo-think.rosetta-ko-tourism-synth-dpo
rosetta-ko-tourism-synth-dpo
Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Tourism Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-tourism-synth-cpt
continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-dpo.rosetta-ko-tourism-synth-dpo-think
rosetta-ko-tourism-synth-dpo-think
Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Tourism Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-tourism-synth-cpt
continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-dpo-think.rosetta-ko-tourism-synth-sft
rosetta-ko-tourism-synth-sft
Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Tourism Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-tourism-synth-cpt
continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-sft.rosetta-ko-tourism-synth-sft-think
rosetta-ko-tourism-synth-sft-think
Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Tourism Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-tourism-synth-cpt
continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-sft-think.rosetta-ko-code-synth-dpo-think
rosetta-ko-code-synth-dpo-think
Korean-native coding data — problems with solutions whose unit tests were actually executed and passed (execution-grounded rejection sampling).
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Code Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-code-synth-sft
supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-code-synth-dpo-think.rosetta-ko-code-synth-rlvr
rosetta-ko-code-synth-rlvr
Korean-native coding data — problems with solutions whose unit tests were actually executed and passed (execution-grounded rejection sampling).
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Code Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-code-synth-sft
supervised fine-tuning (user/assistant… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-code-synth-rlvr.rosetta-ko-heritage-synth-dpo
rosetta-ko-heritage-synth-dpo
Korean cultural-heritage data grounded in public heritage records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Cultural Heritage Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-heritage-synth-cpt
continued-pretraining corpus… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-heritage-synth-dpo.rosetta-ko-heritage-synth-dpo-think
rosetta-ko-heritage-synth-dpo-think
Korean cultural-heritage data grounded in public heritage records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Cultural Heritage Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-heritage-synth-cpt
continued-pretraining… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-heritage-synth-dpo-think.
