datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LFM2.5-KO-CPT-Full-LFMStyle-Raw-20260627
LFM2.5-KO-CPT-Full-LFMStyle-Raw-20260627
Full Korean CPT mix converted to LFM-style text JSONL, about 4B-token training source.
This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow.
Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT
CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL
Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT
SFT GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-CPT-Full-LFMStyle-Raw-20260627.croco-munin-apertus-8b-da-simpo-full-50kcroco-munin-apertus-8b-da-simpo-fullfullmodelaion
