CoolFace
Datasetpublicgated

PoSTMEDIA/rosetta-ko-law-synth-sft

rosetta-ko-law-synth-sft Korean legal-domain data grounded in national statutes — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Law Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-law-synth-cpt continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-law-synth-sft.

sourceHugging Faceapache-2.0updated 12d agoView on Hugging Face
0likes23downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.