CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01iamtarun /python_code_instructions_18k_alpaca Dataset Card for python_code_instructions_18k_alpaca The dataset contains problem descriptions and code in python language. This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here. textquestion-answering10K<n<100K349 likes37k downloads3y agoHugging Face02Nan-Do /code-search-net-python Dataset Card for "code-search-net-python" Dataset Description Homepage: None Repository: https://huggingface.co/datasets/Nan-Do/code-search-net-python Paper: None Leaderboard: None Point of Contact: @Nan-Do Dataset Summary This dataset is the Python portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-python.texttext-generation100K<n<1M30 likes4.2k downloads3y agoHugging Face03jtatman /python-code-dataset-500k Attention: This dataset is a summary and reformat pulled from github code. You should make your own assumptions based on this. In fact, there is another dataset I formed through parsing that addresses several points: out of 500k python related items, most of them are python-ish, not pythonic the majority of the items here contain excessive licensing inclusion of original code the items here are sometimes not even python but have references There's a whole lot of gpl summaries… See the full description on the dataset page: https://huggingface.co/datasets/jtatman/python-code-dataset-500k.texttext-generation100K<n<1M81 likes1.2k downloads3y agoHugging Face04Nan-Do /instructional_code-search-net-python Dataset Card for "instructional_code-search-net-python" Dataset Summary This is an instructional dataset for Python. The dataset contains two different kind of tasks: Given a piece of code generate a description of what it does. Given a description generate a piece of code that fulfils the description. Languages The dataset is in English. Data Splits There are no splits. Dataset Creation May of 2023 Curation Rationale This… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/instructional_code-search-net-python.texttext-generation100K<n<1M36 likes972 downloads3y agoHugging Face05AlgorithmicResearchGroup /arxiv_python_research_code Dataset Card for "ArtifactAI/arxiv_python_research_code" Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_python_research_code Dataset Summary AlgorithmicResearchGroup/arxiv_python_research_code contains over 4.13GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs. How to use it from datasets import load_dataset # full dataset (4.13GB of data) ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_python_research_code.tabulartext-generation1M<n<10M4 likes576 downloads2y agoHugging Face06AlgorithmicResearchGroup /arxiv_deep_learning_python_research_code ArXiv Deep Learning Python Research Code A curated corpus of Python source code files extracted from GitHub repositories referenced in ArXiv papers. Contains 391,496 files (1.49 GB) filtered to deep learning frameworks, designed for training and evaluating Code LLMs on research-grade code. Dataset Summary Statistic Value Total files 391,496 Total size 1.49 GB Source repos 34,099 Time span ArXiv inception through July 2023 Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code.tabulartext-generation100K<n<1M11 likes243 downloads5mo agoHugging Face07ronantakizawa /python-code-instructions-japanese Python Code Instructions - Japanese (18K) Dataset Description This dataset contains 18,612 Python programming instruction-response pairs translated to Japanese. It's designed for training language models to understand and generate Python code based on Japanese instructions. Key Features 18,612 entries covering diverse Python programming tasks Japanese instructions and prompts for code generation Original English text preserved for reference Python code… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/python-code-instructions-japanese.texttext-generation10K<n<100K2 likes199 downloads10mo agoHugging Face08ahmetggg /Dr-Zeon-Github-Python-Code-Dataset Luck Spark 1B - High Quality Code Dataset The first quality-scored, star-agnostic code dataset for training 1B MoE code models. Unlike The Stack / CodeParrot that filter by stars, this dataset scores every file by its content (0-10). A 2-star well-documented library scores higher than a 10k-star minified file. Continuously updated by an autonomous bot. Repo: ahmetggg/luck-spark-1b-code-dataset | Bot: github_to_hf_bot.py | License: Permissive only (MIT / Apache-2.0 / BSD /… See the full description on the dataset page: https://huggingface.co/datasets/ahmetggg/Dr-Zeon-Github-Python-Code-Dataset.tabulartext-generation10K<n<100K1 likes142 downloads25d agoHugging Face09jtatman /combined_coder_pythonCombining smaller python code datasets into a larger one. Changed format to system, instruction, output. Built from: dataset1: nickrosh/Evol-Instruct-Code-80k-v1 dataset2: ehartford/dolphin-coder dataset3: iamtarun/python_code_instructions_18k_alpaca dataset4: iamtarun/python_code_instructions_18k_alpaca dataset5: Vezora/Tested-22k-Python-Alpaca dataset6: mlabonne/Evol-Instruct-Python-26k dataset7: KrisPi/PythonTutor-Evol-1k-DPO-GPT4_vs_35 dataset8:… See the full description on the dataset page: https://huggingface.co/datasets/jtatman/combined_coder_python.texttext-generation100K<n<1M5 likes127 downloads2y agoHugging Face10Orion-zhen /github-python-code-fim github python code fim Generated from tomekkorbak/python-github-code, limiting context length to 8192. tabulartext-generation100K<n<1M0 likes124 downloads1y agoHugging Face11Ananda100 /python-clean-codeparrot python-clean-codeparrot A cleaned, deduplicated, Python-only pretraining corpus derived from codeparrot/codeparrot-clean, built as the pretraining data for PocketCoder, a 95.87M-parameter decoder-only code language model. 1,800,000 documents, ~2.96 billion tokens (DeepSeek-Coder tokenizer, vocabulary 32,022). Paper: PocketCoder: What Distillation, SFT, and DPO Each Buy You at 100M Parameters Model: Ananda100/PocketCoder SFT dataset: Ananda100/python-sft-dataset Code:… See the full description on the dataset page: https://huggingface.co/datasets/Ananda100/python-clean-codeparrot.texttext-generation1M<n<10M0 likes115 downloads1mo agoHugging Face12Nan-Do /reason_code-search-net-python Dataset Card for "reason_code-search-net-python" Dataset Summary This dataset is an instructional dataset for Python.The dataset contains five different kind of tasks. Given a Python 3 function: Type 1: Generate a summary explaining what it does. (For example: This function counts the number of objects stored in the jsonl file passed as input.) Type 2: Generate a summary explaining what its input parameters represent ("For example: infile: a file descriptor of a file… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/reason_code-search-net-python.textsummarization100K<n<1M18 likes114 downloads3y agoHugging Face13domofon /python-code-cot-18k python-code-cot-18k COT distilled dataset with 16,565 examples. Source Base: iamtarun/python_code_instructions_18k_alpaca Model: Mistral-7B-Instruct-v0.2-AWQ Format instruction: Task thinking: <think>...</think> reasoning response: Solution texttext-generation10K<n<100K1 likes105 downloads9mo agoHugging Face14Noushad999 /ML-1M-Syntax-Validated-Python-Code ML-1M Syntax-Validated Python Code Dataset Summary ML-1M Syntax-Validated Python Code is a large-scale corpus containing over 1 million machine-learning–oriented Python programs derived from The Stack, a permissively licensed collection of open-source source code. The dataset is constructed through heuristic ML-domain filtering, syntactic validation, and basic safety checks. It is intended to support empirical analysis of real-world ML code, executability and dependency… See the full description on the dataset page: https://huggingface.co/datasets/Noushad999/ML-1M-Syntax-Validated-Python-Code.texttext-generation1M<n<10M0 likes105 downloads8mo agoHugging Face15khemprogrammer /code-search-net-python Dataset Card for "code-search-net-python" Dataset Description Homepage: None Repository: https://huggingface.co/datasets/Nan-Do/code-search-net-python Paper: None Leaderboard: None Point of Contact: @Nan-Do Dataset Summary This dataset is the Python portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the… See the full description on the dataset page: https://huggingface.co/datasets/khemprogrammer/code-search-net-python.texttext-generation100K<n<1M0 likes105 downloads8mo agoHugging Face16AmareshHebbar /leetcode-codegen-python LeetCode Code-Gen Dataset — Python 2522 rows. Given a problem statement, its input/output examples, and a required algorithm/technique, generate a correct Python solution. Part of a 4-language collection built from the same source: see the sibling Python, Java, C++, and JavaScript datasets. Verification Every row was extracted, then executed in a sandboxed subprocess against the problem's own stated examples. Only rows that passed all examples are included -- this… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-python.texttext-generation1K<n<10K1 likes88 downloads3mo agoHugging Face17jtatman /python-github-code-instruct-filtered-5k Dataset Card for "python-github-code-instruct-filtered-5k" This fine dataset tomekkorbak/python-github-code, filtered by scores greater than 0.03. Feedback and additional columns generated through OpenAI and Cohere responses. texttext-generation1K<n<10K7 likes87 downloads2y agoHugging Face18erythropygia /Instruct-Python-Code-Turkish Dataset Card for Instruct-Python-Code-Turkish Language: Turkish Dataset Description The translation was performed using the Google translation model to ensure high-quality, accurate translation. Dataset Details Size: ≈5K Translation tool: Google Translate Data format: Instruct, Output texttext-generation1K<n<10K1 likes61 downloads2y agoHugging Face19mrbesher /python-code-instructions-18k-alpaca-tr Python Code Instructions 18K Alpaca (Turkish) Turkish translation of Python code instruction dataset for code generation tasks. Dataset Details Records: 18,610 Language: Turkish Format: Alpaca-style instruction/input/output Columns Column Description text Formatted training text (instruction + input + code output) instruction Turkish instruction input Optional input/context output Python code solution Example { "text":… See the full description on the dataset page: https://huggingface.co/datasets/mrbesher/python-code-instructions-18k-alpaca-tr.texttext-generation10K<n<100K0 likes43 downloads6mo agoHugging Face20sdfedsge /combined_coder_pythonCombining smaller python code datasets into a larger one. Changed format to system, instruction, output. Built from: dataset1: nickrosh/Evol-Instruct-Code-80k-v1 dataset2: ehartford/dolphin-coder dataset3: iamtarun/python_code_instructions_18k_alpaca dataset4: iamtarun/python_code_instructions_18k_alpaca dataset5: Vezora/Tested-22k-Python-Alpaca dataset6: mlabonne/Evol-Instruct-Python-26k dataset7: KrisPi/PythonTutor-Evol-1k-DPO-GPT4_vs_35 dataset8:… See the full description on the dataset page: https://huggingface.co/datasets/sdfedsge/combined_coder_python.texttext-generation100K<n<1M0 likes43 downloads6d agoHugging Face21ajax9000 /python_code_instructions_18k_alpaca Dataset Card for python_code_instructions_18k_alpaca The dataset contains problem descriptions and code in python language. This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here. textquestion-answering10K<n<100K0 likes41 downloads8d agoHugging Face22lddl16 /python_code Dataset Card for python_code_instructions_18k_alpaca The dataset contains problem descriptions and code in python language. This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here. textquestion-answering10K<n<100K0 likes36 downloads4d agoHugging Face23HachiML /amenokaku-code-instruct-python-mit-450kunishou/amenokaku-code-instructを以下の条件で絞り込んだものです。 MITライセンス (licence: 'MIT') python (source: ['gasyori_100_knocks', 'datascience_100_knocks_python', 'bifi', 'python_for_begginers_solve_50_exercises', 'nlp_100_knocks') source: 'bifi'をランダムに100件に絞り込み texttext-generationn<1K0 likes30 downloads2y agoHugging Face24gcw-ai /python_code_critic_21k Python Code Critic Dataset Overview This dataset is designed for the automation of generating and validating responses to Python programming questions. It contains data points that consist of a Python question (instruction), a generated response (answer) with code snippets and explanations, the result of code execution (execution_result), an evaluative summary (thought), a determination of response appropriateness (action), and, if necessary, an improved answer… See the full description on the dataset page: https://huggingface.co/datasets/gcw-ai/python_code_critic_21k.texttext-generation10K<n<100K0 likes27 downloads2y agoHugging Face25me-aas /python-code-dataset-500k Attention: This dataset is a summary and reformat pulled from github code. You should make your own assumptions based on this. In fact, there is another dataset I formed through parsing that addresses several points: out of 500k python related items, most of them are python-ish, not pythonic the majority of the items here contain excessive licensing inclusion of original code the items here are sometimes not even python but have references There's a whole lot of gpl summaries… See the full description on the dataset page: https://huggingface.co/datasets/me-aas/python-code-dataset-500k.texttext-generation100K<n<1M0 likes26 downloads4mo agoHugging Face26HachiML /amenokaku-code-instruct-python-mitkunishou/amenokaku-code-instructを以下の条件で絞り込んだものです。 MITライセンス (licence: 'MIT') python (source: ['gasyori_100_knocks', 'datascience_100_knocks_python', 'bifi', 'python_for_begginers_solve_50_exercises', 'nlp_100_knocks') texttext-generation1K<n<10K0 likes25 downloads2y agoHugging Face27iamkoder001 /python_code_instructions_18k_alpaca Dataset Card for python_code_instructions_18k_alpaca The dataset contains problem descriptions and code in python language. This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here. textquestion-answering10K<n<100K0 likes24 downloads7mo agoHugging Face28kd13 /PythonCodeChat-v1 PythonCodeChat-v1 PythonCodeChat-v1 is a synthetic multi-turn Python programming dataset designed for supervised fine-tuning of conversational language models. It contains interactive coding dialogues covering Python concepts, debugging, code generation, optimization, refactoring, library usage, error fixing, and step-by-step programming assistance across multiple conversation turns. The dataset is suitable for training Python coding assistants, educational tutors, and… See the full description on the dataset page: https://huggingface.co/datasets/kd13/PythonCodeChat-v1.texttext-generationn<1K0 likes24 downloads3mo agoHugging Face29kd13 /PythonCodeChat-v2-Reasoning PythonCodeChat-v2-Reasoning PythonCodeChat-v2-Reasoning is a synthetic multi-turn Python programming dataset designed for supervised fine-tuning of language models with enhanced reasoning capabilities. It contains interactive coding conversations that emphasize step-by-step problem solving, debugging, code generation, optimization, refactoring, and Python best practices across multiple dialogue turns. The dataset is suitable for training reasoning-based Python coding assistants… See the full description on the dataset page: https://huggingface.co/datasets/kd13/PythonCodeChat-v2-Reasoning.texttext-generation1K<n<10K0 likes24 downloads3mo agoHugging Face30jonathanyly /python-code-dataset-500k Attention: This dataset is a summary and reformat pulled from github code. You should make your own assumptions based on this. In fact, there is another dataset I formed through parsing that addresses several points: out of 500k python related items, most of them are python-ish, not pythonic the majority of the items here contain excessive licensing inclusion of original code the items here are sometimes not even python but have references There's a whole lot of gpl summaries… See the full description on the dataset page: https://huggingface.co/datasets/jonathanyly/python-code-dataset-500k.texttext-generation100K<n<1M1 likes21 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.