CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01code-rag-bench /github-repos-pythonThe Github repository retrieval source for [code-rag-bench], containing all Python files from the entire GitHub dump (in github-repos) text100K<n<1M2 likes3.3k downloads2y agoHugging Face02ajibawa-2023 /Python-Code-23k-ShareGPTThis dataset is in Vicuna/ShareGPT format. There are 23000+ set of conversations. Each set having 2 conversations. Along with the Python code detailed explanation is provided. This dataset was generated using GPT-3.5, GPT-4 etc. text10K<n<100K42 likes1.2k downloads3y agoHugging Face03Lovett01 /Python-Code-LargePython-Code-Large Python-Code-Large is a large-scale corpus of Python source code comprising more than 2 million rows of Python code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the Python ecosystem. By providing a high-volume, language-specific corpus, Python-Code-Large enables systematic experimentation in Python-focused model training, domain adaptation, and downstream… See the full description on the dataset page: https://huggingface.co/datasets/Lovett01/Python-Code-Large.texttext-generation1M<n<10M0 likes721 downloads2mo agoHugging Face04ajibawa-2023 /Python-Code-LargePython-Code-Large Python-Code-Large is a large-scale corpus of Python source code comprising more than 2 million rows of Python code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the Python ecosystem. By providing a high-volume, language-specific corpus, Python-Code-Large enables systematic experimentation in Python-focused model training, domain adaptation, and downstream… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Python-Code-Large.texttext-generation1M<n<10M19 likes660 downloads7mo agoHugging Face05NickIBrody /python-code-instructions-85k Python Code Instructions - 85K Instruction-tuning dataset of Python functions paired with short natural-language instructions derived from repository docstrings. What changed in this release This release keeps the original public rows and format, but makes the dataset easier to use responsibly: exact duplicate rows were removed again using normalized instruction + output hashing deterministic train, validation, and test splits were added the dataset card now documents… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/python-code-instructions-85k.texttext-generation10K<n<100K1 likes308 downloads5mo agoHugging Face06semeru /code-text-python Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-text/python in Semeru CodeXGLUE -- Code-To-Text Task Definition The task is to generate natural language comments for a code, and evaluted by smoothed bleu-4 score. Dataset The dataset we use comes from CodeSearchNet and we filter the dataset as the following: Remove… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-text-python.text100K<n<1M11 likes300 downloads4y agoHugging Face07karths /python_codetextn<1K1 likes189 downloads2y agoHugging Face08bunyaminergen /Stable-Code-Python-SFT Stable Code Python SFT The Stable Code Python SFT dataset is a high-quality synthetic dataset derived from the stabilityai/stable-code-instruct-3b model for the purpose of supervised fine-tuning (SFT). Please refer to the Versioning section for dataset versions. Note: If you would like to contribute to this repository, please read the CONTRIBUTING first. TableofContents Features File Structure Metadata Usage Versioning License TeamContact Reference Citation… See the full description on the dataset page: https://huggingface.co/datasets/bunyaminergen/Stable-Code-Python-SFT.textquestion-answering10K<n<100K2 likes186 downloads1y agoHugging Face09Mir-2002 /python_code_docstring_ast_corpus Overview This dataset contains 34,000+ rows of code-docstring-ast data along with additional metadata. Data was gathered from various Python libraries and frameworks and their publicly available GitHub repos. This dataset was created for the purpose of training the CodeT5+ transformer on AST-enhanced code-to-doc tasks. Sources The dataset was gathered from various GitHub repos sampled from this repo by Vinta. The 26 repos are: matplotlib pytorch cryptography django… See the full description on the dataset page: https://huggingface.co/datasets/Mir-2002/python_code_docstring_ast_corpus.textsummarization10K<n<100K1 likes122 downloads1y agoHugging Face10jeffmeloy /python_documentation_codeDataset created using https://github.com/jeffmeloy/py2dataset using the Python code from the following: https://github.com/ansible/ansible https://github.com/apache/airflow https://github.com/arogozhnikov/einops https://github.com/arviz-devs/arviz https://github.com/astropy/astropy https://github.com/biopython/biopython https://github.com/bjodah/chempy https://github.com/bokeh/bokehhttps://github.com/CalebBell/thermo https://github.com/camDavidsonPilon/lifelines https://github.com/coin-or/pulp… See the full description on the dataset page: https://huggingface.co/datasets/jeffmeloy/python_documentation_code.texttext-generation10K<n<100K1 likes116 downloads2y agoHugging Face11Veri-Code /ReForm-Python2Dafny-Dataset Re:Form Datasets This repository contains the datasets associated with the paper "Re:Form -- Reducing Human Priors in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny". The paper introduces a framework that leverages Reinforcement Learning (RL) within Large Language Models (LLMs) to reduce reliance on human-annotated priors for formal software verification. The datasets provided are integral for training and evaluating models in formal language… See the full description on the dataset page: https://huggingface.co/datasets/Veri-Code/ReForm-Python2Dafny-Dataset.texttext-generation10K<n<100K2 likes100 downloads5mo agoHugging Face12theprint /MultiRoundConvos-Code-JS-HTML-CSS-Pythontext1K<n<10K0 likes83 downloads9mo agoHugging Face13Ramikan-BR /code.evol.instruct.wiz.oss_python.jsontabulartext-generation1K<n<10K0 likes40 downloads2y agoHugging Face14Wanfq /python_codehttps://huggingface.co/datasets/ajibawa-2023/Python-Code-23k-ShareGPT features: coding, single-turn, task length: 22.6k text10K<n<100K3 likes35 downloads3y agoHugging Face15gonglinyuan /code_search_net_python_tokenizedtext100K<n<1M2 likes33 downloads3y agoHugging Face16kalomaze /code-judge-ast-python-15k-it1 code-judge-ast-python-15k-it1 15k examples of Python code + yes/no structural property questions with deterministic AST ground truth. 3 points per example, balanced to 40-60% per property. Source: The Stack v1 (deduplicated Python). text10K<n<100K0 likes26 downloads8mo agoHugging Face17Myashka /SO-Python_QA-filtered-2023-no_code-tanh_scoreSO dataset of pythontag data Question filters: images links code blocks Q_Score > 0 Answer_count > 0 Answers filters: images links code blocks Scores are tanh applied to scaled with AbsMaxScaler to IQR range of Original SO Answers' scores tabularquestion-answering10K<n<100K2 likes23 downloads3y agoHugging Face18Guilherme34 /openai-finetune-python-code-instructionstext10K<n<100K0 likes19 downloads9mo agoHugging Face19MohamedSaeed-dev /python-text-to-codetext10K<n<100K8 likes18 downloads3y agoHugging Face20waltertaya /python-mal-code-dataset Python Code Dataset This dataset contains extracted Python code from various repositories for fine-tuning code-generation models. text1K<n<10K0 likes15 downloads9mo agoHugging Face21neuronpedia-org /python-code-simplified What This Is This is a dataset of "simplified" Python Code from https://huggingface.co/datasets/iamtarun/python_code_instructions_18k_alpaca. Our simplification attempts to remove all comments, and reduce these names/strings to a single letter without impacting its structure/logic. You should use the "minimized" field. The "original" field is the original output from iamtarun/python_code_instructions_18k_alpaca. Why It Exists We used this dataset to generate… See the full description on the dataset page: https://huggingface.co/datasets/neuronpedia-org/python-code-simplified.text10K<n<100K0 likes15 downloads5mo agoHugging Face22phongmt184172 /python_code_version2text10K<n<100K0 likes14 downloads3y agoHugging Face23nadiamaqbool81 /python_code_instructions_510_alpacatextn<1K0 likes14 downloads3y agoHugging Face24pythonist /code_instruction_alpaca_kkmtextn<1K0 likes12 downloads3y agoHugging Face25bharathgp /cb_qa_python_codetext1K<n<10K0 likes9 downloads2y agoHugging Face26frankminors123 /Python-Code-Instructions-7ktext1K<n<10K0 likes8 downloads3y agoHugging Face27superchikkibaby /python-code-trainingtext10K<n<100K0 likes8 downloads1mo agoHugging Face28minakaragoz /Python_Codetext1K<n<10K0 likes6 downloads5mo agoHugging Face29kubabp9 /python_code_instructions_1k_alpacatext10K<n<100K0 likes3 downloads2y agoHugging Face30Joshu66 /python_code_instructions_18k_chatML_chinese2text10K<n<100K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.