datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
python-github-code-instruct-filtered-5k
Dataset Card for "python-github-code-instruct-filtered-5k"
This fine dataset tomekkorbak/python-github-code, filtered by scores greater than 0.03.
Feedback and additional columns generated through OpenAI and Cohere responses.
Instruct-Python-Code-Turkish
Dataset Card for Instruct-Python-Code-Turkish
Language: Turkish
Dataset Description
The translation was performed using the Google translation model to ensure high-quality, accurate translation.
Dataset Details
Size: ≈5K
Translation tool: Google Translate
Data format: Instruct, Output
gene-code-generation-instruct
code-generation-instruct v2
Gate-passed instruction data for code-generation — published when 50 fresh examples cleared the quality bar
Kind: synthetic
Domain: code-generation
Records: 96
Created: 2026-06-20T19:02:16+00:00
SHA-256: a3f6a919356ea6d71f365ea93c8ea06cb7dc19f22cad57210dedb1f327caed90
Pipeline: v2.0.0
Filters: {"min_quality": 0.55, "limit": 1000, "source": null, "backend": "llama", "min_judge": 0.7}
Generated by: Qwen3-4B-Instruct-2507-Q4_K_M.gguf (backend:… See the full description on the dataset page: https://huggingface.co/datasets/Gene829/gene-code-generation-instruct.ru_codefeedback_python_Qwen2.5-Coder-32B-Instruct-GPTQ-Int8_sample
ru_Code-Feedback
Вопросы python Code-Feedback
Решение и unit-test с результатами python исполнения.
Made with Qwen2.5-Coder-32B-Instruct-GPTQ-Int8
ru_eval_status
count
OK
2554
Exception
2337
SyntaxError
518
Timeout
79
aihub-korean-education-instruct-sample
Korean Education Instruction Dataset (Sample)
Note: 이 데이터셋은 전체 데이터셋의 샘플 버전입니다 (카테고리별 최대 1,000건).
개요
AI Hub의 한국어 교육 데이터셋 13종을 sLLM 지시학습(Instruction Tuning)용으로 변환한 데이터셋입니다.
초등학교부터 고등학교까지의 다양한 교육 콘텐츠를 포함합니다.
데이터셋 통계
카테고리
데이터 수
math (수학)
1000
korean (국어)
1000
writing (글쓰기)
1000
career (진로)
1000
curriculum (교과)
1000
tutor (튜터링)
1000
총계
6000
사용 방법
from datasets import load_dataset
# 데이터셋 로드
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/neuralfoundry-coder/aihub-korean-education-instruct-sample.code_instruct_alpaca
Dataset Card for python_code_instructions_18k_alpaca
The dataset contains problem descriptions and code in python language.
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here.
