CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01iamtarun /python_code_instructions_18k_alpaca Dataset Card for python_code_instructions_18k_alpaca The dataset contains problem descriptions and code in python language. This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here. textquestion-answering10K<n<100K349 likes37k downloads3y agoHugging Face02Nan-Do /code-search-net-python Dataset Card for "code-search-net-python" Dataset Description Homepage: None Repository: https://huggingface.co/datasets/Nan-Do/code-search-net-python Paper: None Leaderboard: None Point of Contact: @Nan-Do Dataset Summary This dataset is the Python portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-python.texttext-generation100K<n<1M30 likes4.2k downloads3y agoHugging Face03angie-chen55 /python-github-codetext1M<n<10M50 likes3k downloads4y agoHugging Face04MatrixStudio /Codeforces-Python-Submissions Dataset Card for "Codeforces-Python-Submissions" More Information needed tabular100K<n<1M45 likes1.4k downloads2y agoHugging Face05jtatman /python-code-dataset-500k Attention: This dataset is a summary and reformat pulled from github code. You should make your own assumptions based on this. In fact, there is another dataset I formed through parsing that addresses several points: out of 500k python related items, most of them are python-ish, not pythonic the majority of the items here contain excessive licensing inclusion of original code the items here are sometimes not even python but have references There's a whole lot of gpl summaries… See the full description on the dataset page: https://huggingface.co/datasets/jtatman/python-code-dataset-500k.texttext-generation100K<n<1M81 likes1.2k downloads3y agoHugging Face06AlgorithmicResearchGroup /arxiv_deep_learning_python_research_code_functions_summaries Dataset Card for "AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries" Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries Dataset Summary AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries contains summaries for every python function and class extracted from source code files referenced in ArXiv papers. The… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries.tabular100K<n<1M9 likes1k downloads2y agoHugging Face07Nan-Do /instructional_code-search-net-python Dataset Card for "instructional_code-search-net-python" Dataset Summary This is an instructional dataset for Python. The dataset contains two different kind of tasks: Given a piece of code generate a description of what it does. Given a description generate a piece of code that fulfils the description. Languages The dataset is in English. Data Splits There are no splits. Dataset Creation May of 2023 Curation Rationale This… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/instructional_code-search-net-python.texttext-generation100K<n<1M36 likes972 downloads3y agoHugging Face08zjsd /RedStone-Code-python RedStone Based on the paper "RedStone: Curating General, Code, Math, and QA Data for Large Language Models" and the official GitHub repository, I have replicated the processing of the RedStone-Code (python only) dataset in Redstone. I followed the processing steps outlined in the official repository with minimal modifications. I have not yet used this data for training to verify its quality. The release is under the Redstone's license. If any data within it infringes on your… See the full description on the dataset page: https://huggingface.co/datasets/zjsd/RedStone-Code-python.text1M<n<10M8 likes909 downloads2y agoHugging Face09tomekkorbak /python-github-codetabular100K<n<1M13 likes625 downloads4y agoHugging Face10vikp /python_code_instructions_filtered Dataset Card for "code_filtered" This includes data from xlcost, evol instruct, code alpaca, code instructions, and code search net. Data is filtered based on quality and learning value. text100K<n<1M5 likes601 downloads3y agoHugging Face11AlgorithmicResearchGroup /arxiv_python_research_code Dataset Card for "ArtifactAI/arxiv_python_research_code" Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_python_research_code Dataset Summary AlgorithmicResearchGroup/arxiv_python_research_code contains over 4.13GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs. How to use it from datasets import load_dataset # full dataset (4.13GB of data) ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_python_research_code.tabulartext-generation1M<n<10M4 likes576 downloads2y agoHugging Face12theothertom /codeparrot-python-only 🐍 CodeParrot Python Only This dataset contains Python-only source code extracted from the larger CodeParrot corpus. It includes high-quality .py files filtered from public GitHub repositories and curated for use in training large language models (LLMs) on Python code generation tasks. 📦 Dataset Summary ✅ Filtered to include only Python code 🧹 Cleaned to remove non-source content (e.g., binaries, notebooks, scripts with mixed languages) 🧠 Ideal for training or… See the full description on the dataset page: https://huggingface.co/datasets/theothertom/codeparrot-python-only.text10K<n<100K1 likes554 downloads1y agoHugging Face13allenai /rlvr-code-data-python-r1-format-filteredtabular10K<n<100K4 likes378 downloads1y agoHugging Face14evanellis /evanellis_Codeforces-Python-Submissions_correct_with_h_a_k_prob_0.5_with_null_and_rejected_f_ztabular10K<n<100K0 likes342 downloads2y agoHugging Face15UUUUUUZ /python-code-100-copiedtext1M<n<10M0 likes284 downloads2y agoHugging Face16ammarnasr /Python-React-Code-Datasettabular1K<n<10K2 likes275 downloads3y agoHugging Face17AlgorithmicResearchGroup /arxiv_deep_learning_python_research_code ArXiv Deep Learning Python Research Code A curated corpus of Python source code files extracted from GitHub repositories referenced in ArXiv papers. Contains 391,496 files (1.49 GB) filtered to deep learning frameworks, designed for training and evaluating Code LLMs on research-grade code. Dataset Summary Statistic Value Total files 391,496 Total size 1.49 GB Source repos 34,099 Time span ArXiv inception through July 2023 Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code.tabulartext-generation100K<n<1M11 likes243 downloads5mo agoHugging Face18CM /codexglue_code2text_python Dataset Card for "codexglue_code2text_python" More Information needed text100K<n<1M8 likes219 downloads3y agoHugging Face19evanellis /Codeforces-Python_Submissions_reformatted_deduped_llama3.3_nomic_ftabular10K<n<100K0 likes206 downloads1y agoHugging Face20windchimeran /codenet_pythonThis is dataset is extracted from CodeNet, python only. I merged the data into one single table, including metadata, problem description, test input output. small: accepted status only big: all status, including accepted tabular1M<n<10M1 likes202 downloads1y agoHugging Face21infinityofspace /python_codestyles-single-500 Dataset Card for "python_codestyles-single-500" This dataset contains negative and positive examples with python code of compliance with a code style. A positive example represents compliance with the code style (label is 1). Each example is composed of two components, the first component consists of a code that either conforms to the code style or violates it and the second component corresponding to an example code that already conforms to a code style. In total, the dataset… See the full description on the dataset page: https://huggingface.co/datasets/infinityofspace/python_codestyles-single-500.tabular100K<n<1M0 likes200 downloads3y agoHugging Face22ronantakizawa /python-code-instructions-japanese Python Code Instructions - Japanese (18K) Dataset Description This dataset contains 18,612 Python programming instruction-response pairs translated to Japanese. It's designed for training language models to understand and generate Python code based on Japanese instructions. Key Features 18,612 entries covering diverse Python programming tasks Japanese instructions and prompts for code generation Original English text preserved for reference Python code… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/python-code-instructions-japanese.texttext-generation10K<n<100K2 likes199 downloads10mo agoHugging Face23ZHENGRAN /cross_code_eval_pythontext1K<n<10K3 likes171 downloads2y agoHugging Face24kejian /codesearchnet-python-raw-457ktext100K<n<1M7 likes166 downloads4y agoHugging Face25kejian /codesearchnet-python-rawtext100K<n<1M2 likes146 downloads4y agoHugging Face26ahmetggg /Dr-Zeon-Github-Python-Code-Dataset Luck Spark 1B - High Quality Code Dataset The first quality-scored, star-agnostic code dataset for training 1B MoE code models. Unlike The Stack / CodeParrot that filter by stars, this dataset scores every file by its content (0-10). A 2-star well-documented library scores higher than a 10k-star minified file. Continuously updated by an autonomous bot. Repo: ahmetggg/luck-spark-1b-code-dataset | Bot: github_to_hf_bot.py | License: Permissive only (MIT / Apache-2.0 / BSD /… See the full description on the dataset page: https://huggingface.co/datasets/ahmetggg/Dr-Zeon-Github-Python-Code-Dataset.tabulartext-generation10K<n<100K1 likes142 downloads25d agoHugging Face27natolambert /rlvr-code-data-python-r1tabular10K<n<100K2 likes141 downloads1y agoHugging Face28AlgorithmicResearchGroup /arxiv_python_research_code_summaries Dataset Card for "ArtifactAI/arxiv_python_research_code_summaries" Dataset Description https://huggingface.co/datasets/ArtifactAI/arxiv_python_research_code_summaries Dataset Summary ArtifactAI/arxiv_deep_learning_python_research_code contains summaries for every python function and class extracted from source code files referenced in ArXiv papers. The dataset serves as a curated dataset for Code LLMs. How to use it from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_python_research_code_summaries.text100K<n<1M0 likes140 downloads2y agoHugging Face29infinityofspace /python_codestyles-mixed1-1k Dataset Card for "python_codestyles-mixed1-1k" This dataset contains negative and positive examples with python code of compliance with a code style. A positive example represents compliance with the code style (label is 1). Each example is composed of two components, the first component consists of a code that either conforms to the code style or violates it and the second component corresponding to an example code that already conforms to a code style. The dataset combines both… See the full description on the dataset page: https://huggingface.co/datasets/infinityofspace/python_codestyles-mixed1-1k.tabular100K<n<1M0 likes131 downloads3y agoHugging Face30LimYeri /CodeFeedback-Filtered-Instruction-Pythontext100K<n<1M0 likes129 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.