datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
python_code_instructions_18k_alpaca
Dataset Card for python_code_instructions_18k_alpaca
The dataset contains problem descriptions and code in python language.
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here.
code-search-net-python
Dataset Card for "code-search-net-python"
Dataset Description
Homepage: None
Repository: https://huggingface.co/datasets/Nan-Do/code-search-net-python
Paper: None
Leaderboard: None
Point of Contact: @Nan-Do
Dataset Summary
This dataset is the Python portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-python.python-code-dataset-500k
Attention: This dataset is a summary and reformat pulled from github code.
You should make your own assumptions based on this.
In fact, there is another dataset I formed through parsing that addresses several points:
out of 500k python related items, most of them are python-ish, not pythonic
the majority of the items here contain excessive licensing inclusion of original code
the items here are sometimes not even python but have references
There's a whole lot of gpl summaries… See the full description on the dataset page: https://huggingface.co/datasets/jtatman/python-code-dataset-500k.instructional_code-search-net-python
Dataset Card for "instructional_code-search-net-python"
Dataset Summary
This is an instructional dataset for Python.
The dataset contains two different kind of tasks:
Given a piece of code generate a description of what it does.
Given a description generate a piece of code that fulfils the description.
Languages
The dataset is in English.
Data Splits
There are no splits.
Dataset Creation
May of 2023
Curation Rationale
This… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/instructional_code-search-net-python.arxiv_python_research_code
Dataset Card for "ArtifactAI/arxiv_python_research_code"
Dataset Description
https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_python_research_code
Dataset Summary
AlgorithmicResearchGroup/arxiv_python_research_code contains over 4.13GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs.
How to use it
from datasets import load_dataset
# full dataset (4.13GB of data)
ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_python_research_code.arxiv_deep_learning_python_research_code
ArXiv Deep Learning Python Research Code
A curated corpus of Python source code files extracted from GitHub repositories referenced in ArXiv papers. Contains 391,496 files (1.49 GB) filtered to deep learning frameworks, designed for training and evaluating Code LLMs on research-grade code.
Dataset Summary
Statistic
Value
Total files
391,496
Total size
1.49 GB
Source repos
34,099
Time span
ArXiv inception through July 2023
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code.python-code-instructions-japanese
Python Code Instructions - Japanese (18K)
Dataset Description
This dataset contains 18,612 Python programming instruction-response pairs translated to Japanese. It's designed for training language models to understand and generate Python code based on Japanese instructions.
Key Features
18,612 entries covering diverse Python programming tasks
Japanese instructions and prompts for code generation
Original English text preserved for reference
Python code… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/python-code-instructions-japanese.Dr-Zeon-Github-Python-Code-Dataset
Luck Spark 1B - High Quality Code Dataset
The first quality-scored, star-agnostic code dataset for training 1B MoE code models.
Unlike The Stack / CodeParrot that filter by stars, this dataset scores every file by its content (0-10). A 2-star well-documented library scores higher than a 10k-star minified file. Continuously updated by an autonomous bot.
Repo: ahmetggg/luck-spark-1b-code-dataset | Bot: github_to_hf_bot.py | License: Permissive only (MIT / Apache-2.0 / BSD /… See the full description on the dataset page: https://huggingface.co/datasets/ahmetggg/Dr-Zeon-Github-Python-Code-Dataset.combined_coder_pythonCombining smaller python code datasets into a larger one.
Changed format to system, instruction, output.
Built from:
dataset1: nickrosh/Evol-Instruct-Code-80k-v1
dataset2: ehartford/dolphin-coder
dataset3: iamtarun/python_code_instructions_18k_alpaca
dataset4: iamtarun/python_code_instructions_18k_alpaca
dataset5: Vezora/Tested-22k-Python-Alpaca
dataset6: mlabonne/Evol-Instruct-Python-26k
dataset7: KrisPi/PythonTutor-Evol-1k-DPO-GPT4_vs_35
dataset8:… See the full description on the dataset page: https://huggingface.co/datasets/jtatman/combined_coder_python.github-python-code-fim
github python code fim
Generated from tomekkorbak/python-github-code, limiting context length to 8192.
python-clean-codeparrot
python-clean-codeparrot
A cleaned, deduplicated, Python-only pretraining corpus derived from
codeparrot/codeparrot-clean,
built as the pretraining data for PocketCoder, a 95.87M-parameter decoder-only
code language model. 1,800,000 documents, ~2.96 billion tokens
(DeepSeek-Coder tokenizer, vocabulary 32,022).
Paper: PocketCoder: What Distillation, SFT, and DPO Each Buy You at 100M Parameters
Model: Ananda100/PocketCoder
SFT dataset: Ananda100/python-sft-dataset
Code:… See the full description on the dataset page: https://huggingface.co/datasets/Ananda100/python-clean-codeparrot.reason_code-search-net-python
Dataset Card for "reason_code-search-net-python"
Dataset Summary
This dataset is an instructional dataset for Python.The dataset contains five different kind of tasks.
Given a Python 3 function:
Type 1: Generate a summary explaining what it does. (For example: This function counts the number of objects stored in the jsonl file passed as input.)
Type 2: Generate a summary explaining what its input parameters represent ("For example: infile: a file descriptor of a file… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/reason_code-search-net-python.python-code-cot-18k
python-code-cot-18k
COT distilled dataset with 16,565 examples.
Source
Base: iamtarun/python_code_instructions_18k_alpaca
Model: Mistral-7B-Instruct-v0.2-AWQ
Format
instruction: Task
thinking: <think>...</think> reasoning
response: Solution
ML-1M-Syntax-Validated-Python-Code
ML-1M Syntax-Validated Python Code
Dataset Summary
ML-1M Syntax-Validated Python Code is a large-scale corpus containing over 1 million machine-learning–oriented Python programs derived from The Stack, a permissively licensed collection of open-source source code.
The dataset is constructed through heuristic ML-domain filtering, syntactic validation, and basic safety checks. It is intended to support empirical analysis of real-world ML code, executability and dependency… See the full description on the dataset page: https://huggingface.co/datasets/Noushad999/ML-1M-Syntax-Validated-Python-Code.code-search-net-python
Dataset Card for "code-search-net-python"
Dataset Description
Homepage: None
Repository: https://huggingface.co/datasets/Nan-Do/code-search-net-python
Paper: None
Leaderboard: None
Point of Contact: @Nan-Do
Dataset Summary
This dataset is the Python portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the… See the full description on the dataset page: https://huggingface.co/datasets/khemprogrammer/code-search-net-python.leetcode-codegen-python
LeetCode Code-Gen Dataset — Python
2522 rows. Given a problem statement, its input/output examples, and a
required algorithm/technique, generate a correct Python solution.
Part of a 4-language collection built from the same source: see the sibling
Python,
Java,
C++, and
JavaScript
datasets.
Verification
Every row was extracted, then executed in a sandboxed subprocess against the problem's own stated examples. Only rows that passed all examples are included -- this… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-python.python-github-code-instruct-filtered-5k
Dataset Card for "python-github-code-instruct-filtered-5k"
This fine dataset tomekkorbak/python-github-code, filtered by scores greater than 0.03.
Feedback and additional columns generated through OpenAI and Cohere responses.
Instruct-Python-Code-Turkish
Dataset Card for Instruct-Python-Code-Turkish
Language: Turkish
Dataset Description
The translation was performed using the Google translation model to ensure high-quality, accurate translation.
Dataset Details
Size: ≈5K
Translation tool: Google Translate
Data format: Instruct, Output
python-code-instructions-18k-alpaca-tr
Python Code Instructions 18K Alpaca (Turkish)
Turkish translation of Python code instruction dataset for code generation tasks.
Dataset Details
Records: 18,610
Language: Turkish
Format: Alpaca-style instruction/input/output
Columns
Column
Description
text
Formatted training text (instruction + input + code output)
instruction
Turkish instruction
input
Optional input/context
output
Python code solution
Example
{
"text":… See the full description on the dataset page: https://huggingface.co/datasets/mrbesher/python-code-instructions-18k-alpaca-tr.combined_coder_pythonCombining smaller python code datasets into a larger one.
Changed format to system, instruction, output.
Built from:
dataset1: nickrosh/Evol-Instruct-Code-80k-v1
dataset2: ehartford/dolphin-coder
dataset3: iamtarun/python_code_instructions_18k_alpaca
dataset4: iamtarun/python_code_instructions_18k_alpaca
dataset5: Vezora/Tested-22k-Python-Alpaca
dataset6: mlabonne/Evol-Instruct-Python-26k
dataset7: KrisPi/PythonTutor-Evol-1k-DPO-GPT4_vs_35
dataset8:… See the full description on the dataset page: https://huggingface.co/datasets/sdfedsge/combined_coder_python.python_code_instructions_18k_alpaca
Dataset Card for python_code_instructions_18k_alpaca
The dataset contains problem descriptions and code in python language.
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here.
python_code
Dataset Card for python_code_instructions_18k_alpaca
The dataset contains problem descriptions and code in python language.
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here.
amenokaku-code-instruct-python-mit-450kunishou/amenokaku-code-instructを以下の条件で絞り込んだものです。
MITライセンス (licence: 'MIT')
python (source: ['gasyori_100_knocks', 'datascience_100_knocks_python', 'bifi', 'python_for_begginers_solve_50_exercises', 'nlp_100_knocks')
source: 'bifi'をランダムに100件に絞り込み
python_code_critic_21k
Python Code Critic Dataset
Overview
This dataset is designed for the automation of generating and validating responses to Python programming questions. It contains data points that consist of a Python question (instruction), a generated response (answer) with code snippets and explanations, the result of code execution (execution_result), an evaluative summary (thought), a determination of response appropriateness (action), and, if necessary, an improved answer… See the full description on the dataset page: https://huggingface.co/datasets/gcw-ai/python_code_critic_21k.python-code-dataset-500k
Attention: This dataset is a summary and reformat pulled from github code.
You should make your own assumptions based on this.
In fact, there is another dataset I formed through parsing that addresses several points:
out of 500k python related items, most of them are python-ish, not pythonic
the majority of the items here contain excessive licensing inclusion of original code
the items here are sometimes not even python but have references
There's a whole lot of gpl summaries… See the full description on the dataset page: https://huggingface.co/datasets/me-aas/python-code-dataset-500k.amenokaku-code-instruct-python-mitkunishou/amenokaku-code-instructを以下の条件で絞り込んだものです。
MITライセンス (licence: 'MIT')
python (source: ['gasyori_100_knocks', 'datascience_100_knocks_python', 'bifi', 'python_for_begginers_solve_50_exercises', 'nlp_100_knocks')
python_code_instructions_18k_alpaca
Dataset Card for python_code_instructions_18k_alpaca
The dataset contains problem descriptions and code in python language.
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here.
PythonCodeChat-v1
PythonCodeChat-v1
PythonCodeChat-v1 is a synthetic multi-turn Python programming dataset designed for supervised fine-tuning of conversational language models. It contains interactive coding dialogues covering Python concepts, debugging, code generation, optimization, refactoring, library usage, error fixing, and step-by-step programming assistance across multiple conversation turns. The dataset is suitable for training Python coding assistants, educational tutors, and… See the full description on the dataset page: https://huggingface.co/datasets/kd13/PythonCodeChat-v1.PythonCodeChat-v2-Reasoning
PythonCodeChat-v2-Reasoning
PythonCodeChat-v2-Reasoning is a synthetic multi-turn Python programming dataset designed for supervised fine-tuning of language models with enhanced reasoning capabilities. It contains interactive coding conversations that emphasize step-by-step problem solving, debugging, code generation, optimization, refactoring, and Python best practices across multiple dialogue turns. The dataset is suitable for training reasoning-based Python coding assistants… See the full description on the dataset page: https://huggingface.co/datasets/kd13/PythonCodeChat-v2-Reasoning.python-code-dataset-500k
Attention: This dataset is a summary and reformat pulled from github code.
You should make your own assumptions based on this.
In fact, there is another dataset I formed through parsing that addresses several points:
out of 500k python related items, most of them are python-ish, not pythonic
the majority of the items here contain excessive licensing inclusion of original code
the items here are sometimes not even python but have references
There's a whole lot of gpl summaries… See the full description on the dataset page: https://huggingface.co/datasets/jonathanyly/python-code-dataset-500k.
