CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mike-ravkine /rosettacode-parsed Data Origins Original dataset: https://huggingface.co/datasets/jondurbin/rosettacode-raw/ Cleaner code: https://github.com/the-crypt-keeper/rosettacode-parser Data Fields Field Type Description title string problem title task string problem description language string solution language/variant soulution string solution source code Languages One .jsonl is provided per language group, the sublanguage field in the data denotes the… See the full description on the dataset page: https://huggingface.co/datasets/mike-ravkine/rosettacode-parsed.texttext-generation1K<n<10K12 likes185 downloads3y agoHugging Face02namanbnsl /rosettabench-150-stratified-compressed About Dataset Why? Used for RosettaBench (A contamination-free benchmark for measuring learning, not memorization). Preparation Source: LiveCodeBench (release_v5, AtCoder platform only), accessed via sam-paech/livecodebench-code_generation_lite on Hugging Face. AtCoder problems were selected for platform consistency and STDIN/STDOUT format compatibility. Problems with fewer than 3 test cases were excluded, leaving a pool of 342 problems. Sampling: 150… See the full description on the dataset page: https://huggingface.co/datasets/namanbnsl/rosettabench-150-stratified-compressed.texttext-generationn<1K2 likes38 downloads5mo agoHugging Face03PoSTMEDIA /rosetta-ko-math-synth-sftgated rosetta-ko-math-synth-sft Korean-native mathematics data — problems with step-by-step Korean solutions, answer-verified by Math-Verify (LaTeX \boxed{} + symbolic equivalence). Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Math Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-math-synth-sft (this repo) supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-math-synth-sft.texttext-generation100K<n<1M0 likes28 downloads12d agoHugging Face04PoSTMEDIA /rosetta-ko-math-synth-sft-thinkgated rosetta-ko-math-synth-sft-think Korean-native mathematics data — problems with step-by-step Korean solutions, answer-verified by Math-Verify (LaTeX \boxed{} + symbolic equivalence). Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Math Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-math-synth-sft supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-math-synth-sft-think.texttext-generation100K<n<1M0 likes26 downloads12d agoHugging Face05PoSTMEDIA /rosetta-ko-chat-synth-sft-thinkgated rosetta-ko-chat-synth-sft-think Korean-native general-purpose conversational data — diverse everyday instructions with natural Korean answers. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Chat Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-chat-synth-sft supervised fine-tuning (user/assistant messages)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-chat-synth-sft-think.texttext-generation100K<n<1M0 likes25 downloads12d agoHugging Face06PoSTMEDIA /rosetta-ko-chat-synth-sftgated rosetta-ko-chat-synth-sft Korean-native general-purpose conversational data — diverse everyday instructions with natural Korean answers. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Chat Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-chat-synth-sft (this repo) supervised fine-tuning (user/assistant messages)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-chat-synth-sft.texttext-generation100K<n<1M0 likes24 downloads12d agoHugging Face07PoSTMEDIA /rosetta-ko-chat-synth-rlvrgated rosetta-ko-chat-synth-rlvr Korean-native general-purpose conversational data — diverse everyday instructions with natural Korean answers. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Chat Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-chat-synth-sft supervised fine-tuning (user/assistant messages)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-chat-synth-rlvr.texttext-generation100K<n<1M0 likes23 downloads12d agoHugging Face08PoSTMEDIA /rosetta-ko-heritage-synth-cptgated rosetta-ko-heritage-synth-cpt Korean cultural-heritage data grounded in public heritage records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Cultural Heritage Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-heritage-synth-cpt (this repo)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-heritage-synth-cpt.text-generation100K<n<1M0 likes23 downloads12d agoHugging Face09PoSTMEDIA /rosetta-ko-law-synth-sftgated rosetta-ko-law-synth-sft Korean legal-domain data grounded in national statutes — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Law Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-law-synth-cpt continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-law-synth-sft.texttext-generation100K<n<1M0 likes23 downloads12d agoHugging Face10PoSTMEDIA /rosetta-ko-math-synth-dpogated rosetta-ko-math-synth-dpo Korean-native mathematics data — problems with step-by-step Korean solutions, answer-verified by Math-Verify (LaTeX \boxed{} + symbolic equivalence). Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Math Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-math-synth-sft supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-math-synth-dpo.text-generation10K<n<100K0 likes23 downloads12d agoHugging Face11PoSTMEDIA /rosetta-ko-math-synth-rlvrgated rosetta-ko-math-synth-rlvr Korean-native mathematics data — problems with step-by-step Korean solutions, answer-verified by Math-Verify (LaTeX \boxed{} + symbolic equivalence). Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Math Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-math-synth-sft supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-math-synth-rlvr.text-generation100K<n<1M0 likes23 downloads12d agoHugging Face12PoSTMEDIA /rosetta-ko-instruction-following-synth-rlvrgated rosetta-ko-instruction-following-synth-rlvr Korean-native instruction-following data — instructions with verifiable constraints (length, format, keywords, JSON, ...) checked by programmatic verifiers. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Instruction-Following Suite Sibling datasets from the same pipeline (each a separate repo): repo format… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-instruction-following-synth-rlvr.texttext-generation100K<n<1M0 likes22 downloads12d agoHugging Face13PoSTMEDIA /rosetta-ko-instruction-following-synth-sft-thinkgated rosetta-ko-instruction-following-synth-sft-think Korean-native instruction-following data — instructions with verifiable constraints (length, format, keywords, JSON, ...) checked by programmatic verifiers. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Instruction-Following Suite Sibling datasets from the same pipeline (each a separate repo): repo format… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-instruction-following-synth-sft-think.texttext-generation10K<n<100K0 likes22 downloads12d agoHugging Face14PoSTMEDIA /rosetta-ko-law-synth-cptgated rosetta-ko-law-synth-cpt Korean legal-domain data grounded in national statutes — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Law Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-law-synth-cpt (this repo) continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-law-synth-cpt.texttext-generation100K<n<1M0 likes22 downloads12d agoHugging Face15PoSTMEDIA /rosetta-ko-law-synth-rlvrgated rosetta-ko-law-synth-rlvr Korean legal-domain data grounded in national statutes — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Law Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-law-synth-cpt continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-law-synth-rlvr.texttext-generation100K<n<1M0 likes22 downloads12d agoHugging Face16PoSTMEDIA /rosetta-ko-tourism-synth-cptgated rosetta-ko-tourism-synth-cpt Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Tourism Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-tourism-synth-cpt (this repo) continued-pretraining corpus (plain… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-cpt.texttext-generation10K<n<100K0 likes22 downloads12d agoHugging Face17PoSTMEDIA /rosetta-ko-tourism-synth-rlvrgated rosetta-ko-tourism-synth-rlvr Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Tourism Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-tourism-synth-cpt continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-rlvr.texttext-generation10K<n<100K0 likes22 downloads12d agoHugging Face18PoSTMEDIA /rosetta-ko-code-synth-sftgated rosetta-ko-code-synth-sft Korean-native coding data — problems with solutions whose unit tests were actually executed and passed (execution-grounded rejection sampling). Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Code Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-code-synth-sft (this repo) supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-code-synth-sft.text-generation100K<n<1M0 likes21 downloads12d agoHugging Face19PoSTMEDIA /rosetta-ko-heritage-synth-rlvrgated rosetta-ko-heritage-synth-rlvr Korean cultural-heritage data grounded in public heritage records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Cultural Heritage Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-heritage-synth-cpt continued-pretraining corpus… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-heritage-synth-rlvr.texttext-generation10K<n<100K0 likes21 downloads12d agoHugging Face20PoSTMEDIA /rosetta-ko-heritage-synth-sft-thinkgated rosetta-ko-heritage-synth-sft-think Korean cultural-heritage data grounded in public heritage records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Cultural Heritage Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-heritage-synth-cpt continued-pretraining… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-heritage-synth-sft-think.texttext-generation100K<n<1M0 likes21 downloads12d agoHugging Face21PoSTMEDIA /rosetta-ko-law-synth-sft-thinkgated rosetta-ko-law-synth-sft-think Korean legal-domain data grounded in national statutes — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Law Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-law-synth-cpt continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-law-synth-sft-think.texttext-generation100K<n<1M0 likes21 downloads12d agoHugging Face22PoSTMEDIA /rosetta-ko-math-synth-dpo-thinkgated rosetta-ko-math-synth-dpo-think Korean-native mathematics data — problems with step-by-step Korean solutions, answer-verified by Math-Verify (LaTeX \boxed{} + symbolic equivalence). Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Math Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-math-synth-sft supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-math-synth-dpo-think.text-generation10K<n<100K0 likes21 downloads12d agoHugging Face23PoSTMEDIA /rosetta-ko-tourism-synth-dpogated rosetta-ko-tourism-synth-dpo Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Tourism Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-tourism-synth-cpt continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-dpo.texttext-generation1K<n<10K0 likes21 downloads12d agoHugging Face24PoSTMEDIA /rosetta-ko-tourism-synth-dpo-thinkgated rosetta-ko-tourism-synth-dpo-think Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Tourism Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-tourism-synth-cpt continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-dpo-think.texttext-generation1K<n<10K0 likes21 downloads12d agoHugging Face25PoSTMEDIA /rosetta-ko-tourism-synth-sftgated rosetta-ko-tourism-synth-sft Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Tourism Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-tourism-synth-cpt continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-sft.texttext-generation10K<n<100K0 likes21 downloads12d agoHugging Face26PoSTMEDIA /rosetta-ko-tourism-synth-sft-thinkgated rosetta-ko-tourism-synth-sft-think Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Tourism Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-tourism-synth-cpt continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-sft-think.texttext-generation10K<n<100K0 likes21 downloads12d agoHugging Face27PoSTMEDIA /rosetta-ko-code-synth-dpo-thinkgated rosetta-ko-code-synth-dpo-think Korean-native coding data — problems with solutions whose unit tests were actually executed and passed (execution-grounded rejection sampling). Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Code Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-code-synth-sft supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-code-synth-dpo-think.texttext-generation10K<n<100K0 likes20 downloads12d agoHugging Face28PoSTMEDIA /rosetta-ko-code-synth-rlvrgated rosetta-ko-code-synth-rlvr Korean-native coding data — problems with solutions whose unit tests were actually executed and passed (execution-grounded rejection sampling). Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Code Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-code-synth-sft supervised fine-tuning (user/assistant… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-code-synth-rlvr.texttext-generation100K<n<1M0 likes20 downloads12d agoHugging Face29PoSTMEDIA /rosetta-ko-heritage-synth-dpogated rosetta-ko-heritage-synth-dpo Korean cultural-heritage data grounded in public heritage records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Cultural Heritage Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-heritage-synth-cpt continued-pretraining corpus… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-heritage-synth-dpo.texttext-generation10K<n<100K0 likes20 downloads12d agoHugging Face30PoSTMEDIA /rosetta-ko-heritage-synth-dpo-thinkgated rosetta-ko-heritage-synth-dpo-think Korean cultural-heritage data grounded in public heritage records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Cultural Heritage Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-heritage-synth-cpt continued-pretraining… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-heritage-synth-dpo-think.texttext-generation10K<n<100K0 likes20 downloads12d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.