datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
1c_github1C_Forums
1C_forums: Dataset based parsed data from two of the most popular forums for 1c
This dataset made of the parsed data from two forums for coders at the 1C languages:
Infostart - all threads presents as like think section for model, and marked as the best result for queestion as final answer
Fastcode - Only templates
Dataset Overview
All rows prepared as useful columns. All text prepared as markdown text, and code 1c looks as like:
\`\`\`1c
"ВЫБРАТЬ
|… See the full description on the dataset page: https://huggingface.co/datasets/arefaste/1C_Forums.1c-enterprise-script-clean-corpus
1C Enterprise Script Clean Corpus
Кратко
Это датасет для обучения моделей работе с кодом на 1C:Enterprise Script.
В репозитории есть два независимых поднабора:
pretrain: чистый корпус реального кода для continued pretraining / domain adaptation
sft_strict: instruction/chat датасет для SFT, построенный поверх очищенного корпуса
Состав
pretrain/train.jsonl
pretrain/validation.jsonl
sft_strict/train.jsonl
sft_strict/validation.jsonl
manifests/repos.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/1c-enterprise-script-clean-corpus.prism-smopPRISM-SMOP
Датасет генераций кода 1С:Предприятие (BSL) с четырёхосевой разметкой качества по метрике SMOP
🤗 Hugging Face ·
Бенчмарк PRISM ·
Лицензия CC BY 4.0 · Метрика SMOP
1225 генераций кода 1С:Предприятие (BSL) от 35 нейросетей на 35 задачах бенчмарка PRISM, размеченных по четырём осям метрики SMOP — синтаксис, смысл, оптимальность, платформа — автоматическим оценщиком L1. Внутри — SFT-подвыборка из 222 решений открытых моделей, прошедших все скрытые тесты.
1225 BSL… See the full description on the dataset page: https://huggingface.co/datasets/genlab-1c/prism-smop.my-distiset-1cff3e34
Dataset Card for my-distiset-1cff3e34
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/hugmah/my-distiset-1cff3e34/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/hugmah/my-distiset-1cff3e34.my-distiset-1ce7d2e1
Dataset Card for my-distiset-1ce7d2e1
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/aturate/my-distiset-1ce7d2e1/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/aturate/my-distiset-1ce7d2e1.
