datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
st-code-datasetst_dataset_distillation_by_qwen2.5.jsonl 是由 stgenerate 蒸馏qwen2.514b生成的
/data/complier_improve 中的是在蒸馏的过程中增加编译器在环`
st_dataset_local.jsonl 是编译验证通过
st_dpo_dataset.jsonl 是交给ai改正后编译验证通过形成DPO
其他的都是各种原因为通过编译
st_dataset_distillation_by_st_coder_clean 这个适合做反面教材
iec61131-3-st-clean-augment
ST-Coder: Multi-Source IEC 61131-3 Dataset
This dataset is designed for training Large Language Models (LLMs) to generate and analyze Structured Text (ST) code according to the IEC 61131-3 industrial standard.
📊 Dataset Subsets
This repository provides multiple configurations based on the source and processing method:
1. usecomplier
Source: Local golden datasets processed via the AST-Augmentation Factory.
Content: High-quality, syntactically correct ST code… See the full description on the dataset page: https://huggingface.co/datasets/RnniaSnow/iec61131-3-st-clean-augment.RNNbC3laQHaS5vEu
