CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01simmo /python-fim Python Stack | Fill-in-the-Middle This is a conversion or adaptation of The Stack to a python FIM task. The example column is B64 encoded because people like to put special characters in their code that csv files dont like so I encoded the strings before saving them to disk. textfill-mask10M<n<100M0 likes4.5k downloads2y agoHugging Face02MCES10-Software /Python-Code-Solutions Python Code Solutions Features 1000k of Python Code Solutions for Text Generation and Question Answering Python Coding Problems labelled by topic and difficulty Recommendations Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models. Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic} textquestion-answering10K<n<100K0 likes598 downloads1y agoHugging Face03espejelomar /code_search_net_python_10000_examplestext10K<n<100K14 likes466 downloads5y agoHugging Face04AmanPriyanshu /random-python-github-repositories random-python-github-repositories A collection of 1650 open-source Python GitHub repositories, packaged as zipped archives alongside a metadata CSV. Intended as a seed dataset for code retrieval, context engineering, and SWE-bench-style dataset construction tasks. All repos contain 250+ .py files. Contents repos_meta_data.csv — metadata for each repo (owner, repo_name, stars, license, py_file_count, alpha_hash) repos-zipped/ — one .zip per repo, named… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/random-python-github-repositories.tabulartext-generation1K<n<10K0 likes235 downloads6mo agoHugging Face05AhmedSSoliman /CodeSearchNet-Pythontext100K<n<1M1 likes192 downloads3y agoHugging Face06Edoh /manim_pythontextn<1K20 likes176 downloads3y agoHugging Face07ananyarn /Algorithm_and_Python_Source_CodeAlgorithm_and_Python_Source_Code This dataset provides different algorithms and their corresponding source code in Python. credits: Source codes given here are taken from "iamtarun/python_code_instructions_18k_alpaca" dataset in Hugging Face. text10K<n<100K11 likes152 downloads3y agoHugging Face08aadajinkya /python_codes_sampletext10K<n<100K2 likes138 downloads3y agoHugging Face09pythonist /PubMedQAtextn<1K0 likes98 downloads4y agoHugging Face10jonaskoenig /ML-Python-Code-Smellstexttext-classificationn<1K1 likes80 downloads2y agoHugging Face11greatdarklord /python_datasettext10K<n<100K1 likes72 downloads3y agoHugging Face12annawleo /python-algorithm-sourcecode Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description This dataset provides algorithms and corresponding Python source code which can be leveraged for any type of code conversion applications. Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/annawleo/python-algorithm-sourcecode.textn<1K0 likes70 downloads3y agoHugging Face13jean1 /45k_python_code_chinese_instruction Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details 中文提示的代码数据集 其中提示部分通过调用GPT-4.0-turbo API翻译成中文 Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/jean1/45k_python_code_chinese_instruction.text10K<n<100K6 likes69 downloads2y agoHugging Face14dylanhogg /awesome-python www.awesomepython.org Hand-picked awesome Python libraries, with an emphasis on data and machine learning 🐍 Dataset used by https://www.awesomepython.org/ license: mit tabulartext-classification1K<n<10K2 likes63 downloads3y agoHugging Face15PythonCreate /Nutrient_Datasettabularn<1K1 likes47 downloads1y agoHugging Face16TacoPrime /errored_pythonThis is a subset of the python dataset provided but Ailurophile on Kaggle. Important:Errors were introduced on purpose to try to test a sort of "specialized masking" in a realistic way. Goal:The goal is to create a specialized agent, and add it to a chain with at least one other agent that generates code, and can hopefully "catch" any errors. Inspiration:When working to generate datasets with other models, I found that even after multiple "passes" errors where still missed. Out of curiosity… See the full description on the dataset page: https://huggingface.co/datasets/TacoPrime/errored_python.texttext-generation10K<n<100K4 likes43 downloads3y agoHugging Face17greatdarklord /python-raw-datasettext10K<n<100K0 likes33 downloads3y agoHugging Face18Bin12345 /fortran-pythontext10K<n<100K3 likes31 downloads3y agoHugging Face19nextpy /CodeExercise-Python-27k-EVOLtext10K<n<100K1 likes31 downloads3y agoHugging Face20guidevit /python_code_summarizationtextn<1K0 likes29 downloads3y agoHugging Face21ASHu2 /docs-python-v1 Dataset Card for Dataset Name This dataset card aims to be a base template for creating python docs from methods. This is formatted from semeru/code-code-galeras-code-completion-from-docstring-3k-deduped Dataset Description Curated by: semeru/code-code-galeras-code-completion-from-docstring-3k-deduped Language(s) (NLP): Python License: [More Information Needed] Dataset Sources [optional] Repository:… See the full description on the dataset page: https://huggingface.co/datasets/ASHu2/docs-python-v1.tabularfeature-extraction1K<n<10K2 likes28 downloads3y agoHugging Face22farahbs /simple_python_descriptiontext1K<n<10K0 likes26 downloads2y agoHugging Face23pythonist /demodtext1K<n<10K0 likes25 downloads3y agoHugging Face24gauravsirola /sas_to_python_base_datasettextn<1K0 likes25 downloads3y agoHugging Face25dbands /pythonMathtext1K<n<10K5 likes23 downloads2y agoHugging Face26jason1966 /PythonforSASUsers_hpindex Home Price Index Housing indexed prices from January 1991 to August 2016 Dataset Info Source: Kaggle Original Size: 0.71 MB Kaggle Downloads: 5,328 Files: 1 Files HPI_master.csv Mirrored from Kaggle tabular10K<n<100K0 likes21 downloads6mo agoHugging Face27abdullahwaseemansari /finance-sales-pythontabular1M<n<10M0 likes20 downloads2mo agoHugging Face28TIGER-Lab /packages_python_filtered SWE-Next: Scalable Real-World Software Engineering Tasks for Agents packages_python_filtered This repository contains packages_python_filtered.csv, the seed repository list used by SWE-Next. The file contains 3,971 Python package / repository entries that serve as the starting point for large-scale repository mining and execution-grounded task synthesis. Each row links a package-oriented seed entry to a GitHub repository and includes lightweight… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/packages_python_filtered.tabular1K<n<10K0 likes19 downloads5mo agoHugging Face29pythonist /nepllmtext10K<n<100K0 likes18 downloads3y agoHugging Face30Mgmgrand420 /code_search_net_python_10000_examplestext10K<n<100K0 likes18 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.