CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01flytech /python-codes-25k License MIT This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks Overview The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects. Dataset Statistics Total Entries: 24,813 Unique Instructions: 24,580 Unique Inputs: 3,666 Unique Outputs: 24,581 Unique Texts: 24,813 Average Tokens per example: 508 Features… See the full description on the dataset page: https://huggingface.co/datasets/flytech/python-codes-25k.texttext-classification10K<n<100K182 likes9.3k downloads2y agoHugging Face02bunyaminergen /Stable-Code-Python-SFT Stable Code Python SFT The Stable Code Python SFT dataset is a high-quality synthetic dataset derived from the stabilityai/stable-code-instruct-3b model for the purpose of supervised fine-tuning (SFT). Please refer to the Versioning section for dataset versions. Note: If you would like to contribute to this repository, please read the CONTRIBUTING first. TableofContents Features File Structure Metadata Usage Versioning License TeamContact Reference Citation… See the full description on the dataset page: https://huggingface.co/datasets/bunyaminergen/Stable-Code-Python-SFT.textquestion-answering10K<n<100K2 likes187 downloads1y agoHugging Face03flytech /llama-python-codes-30k Python Codes - 30k examples, Llama1&2 tokenized dataset Author FlyTech For general guide on how to create, quantize, merge or inference the model and more, visit: hackmd.io/my_first_ai Overview This dataset serves as a rich resource for various Natural Language Processing tasks such as: Question Answering Text Generation Text-to-Text Generation It primarily focuses on instructional tasks in Python, tokenized specifically for the Llama architecture.… See the full description on the dataset page: https://huggingface.co/datasets/flytech/llama-python-codes-30k.textquestion-answering10K<n<100K19 likes91 downloads3y agoHugging Face04creeperdatasets /python_debugging Python Debugging A synthetic instruction-tuning dataset for training AI models to identify and fix bugs in Python code. Dataset Summary Field Value Entries 75 Format input / output pairs Language English Topic Finding and fixing bugs in Python code Synthetic Yes, generated with DeepSeek License MIT Dataset Description Each entry presents a snippet of Python code containing a deliberate bug, along with a corrected version… See the full description on the dataset page: https://huggingface.co/datasets/creeperdatasets/python_debugging.texttext-generationn<1K0 likes91 downloads1mo agoHugging Face05bunyaminergen /Cornstack-Python-V1-Filtered Cornstack Python v1 Filtered The Cornstack Python v1 Filtered dataset is derived from the nomic-ai/cornstack-python-v1 dataset by limiting queries to a maximum of 17 words and restricting the total number of rows to 423259. This dataset is suitable for Python programming education and question-answering applications. Note: If you would like to contribute to this repository, please read the CONTRIBUTING first. TableofContents Features File Structure Metadata Usage… See the full description on the dataset page: https://huggingface.co/datasets/bunyaminergen/Cornstack-Python-V1-Filtered.textquestion-answering100K<n<1M0 likes78 downloads1y agoHugging Face06ILoveBuns /python-mental-execution-traces Python Mental Execution Traces A 12,000-row prompt/completion dataset for evaluating and training language models to mentally execute self-contained Python 3 snippets without running them. Completions provide the expected standard output together with a concise variable trace or explanation. Dataset structure The JSONL file contains two text fields: prompt: a Python mental-execution problem. completion: the expected stdout and concise reasoning or variable trace.… See the full description on the dataset page: https://huggingface.co/datasets/ILoveBuns/python-mental-execution-traces.texttext-generation10K<n<100K0 likes71 downloads2mo agoHugging Face07WithinUsAI /Python_GOD_Coder_Omniforge_AI_12k Python GOD Coder Omniforge AI 12k Creator: Within Us AI A 12,000-row mixed-format Python coding dataset designed as a sharpening corpus for building a small but dangerous Python specialist. This dataset is intentionally focused on the practical behaviors that matter for a modern Python coding model: implementation with tests strict code-only instruction following debugging and repair refactoring for readability and production readiness next-token code completion… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Python_GOD_Coder_Omniforge_AI_12k.texttext-generation10K<n<100K1 likes68 downloads7mo agoHugging Face08xphillyx /python-codes-25k License MIT This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks Overview The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects. Dataset Statistics Total Entries: 24,813 Unique Instructions: 24,580 Unique Inputs: 3,666 Unique Outputs: 24,581 Unique Texts: 24,813 Average Tokens per example: 508 Features… See the full description on the dataset page: https://huggingface.co/datasets/xphillyx/python-codes-25k.texttext-classification10K<n<100K0 likes49 downloads7mo agoHugging Face09MexIvanov /Vezora-Tested-22k-Python-Alpaca-ruA machine translated version of the Vezora/Tested-22k-Python-Alpaca dataset. Consists of code "Filtered Using Vezora's CodeTester" with code-related data and natural language instructions. Released under the same license as the original dataset, provided as is with research intent, use/read at your own risk. textquestion-answering10K<n<100K2 likes47 downloads3y agoHugging Face10nitish26 /python-codes-25k License MIT This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks Overview The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects. Dataset Statistics Total Entries: 24,813 Unique Instructions: 24,580 Unique Inputs: 3,666 Unique Outputs: 24,581 Unique Texts: 24,813 Average Tokens per example: 508 Features… See the full description on the dataset page: https://huggingface.co/datasets/nitish26/python-codes-25k.texttext-classification10K<n<100K1 likes46 downloads9mo agoHugging Face11badaranta /python-codes-25k License MIT This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks Overview The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects. Dataset Statistics Total Entries: 24,813 Unique Instructions: 24,580 Unique Inputs: 3,666 Unique Outputs: 24,581 Unique Texts: 24,813 Average Tokens per example: 508 Features… See the full description on the dataset page: https://huggingface.co/datasets/badaranta/python-codes-25k.texttext-classification10K<n<100K0 likes46 downloads8mo agoHugging Face12bernabepuente /python-instruction-dataset Python Developer Instruction Dataset High-quality instruction-response pairs covering Python development best practices, async programming, decorators, type hints, and data manipulation with Pandas. Dataset Details Dataset Description This is a high-quality instruction-tuning dataset focused on Python topics. Each entry includes: A clear instruction/question Optional input context A detailed response/solution Chain-of-thought reasoning process Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/python-instruction-dataset.texttext-generationn<1K0 likes45 downloads5mo agoHugging Face13Riswan-BluBridge /python-codes-25k License MIT This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks Overview The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects. Dataset Statistics Total Entries: 24,813 Unique Instructions: 24,580 Unique Inputs: 3,666 Unique Outputs: 24,581 Unique Texts: 24,813 Average Tokens per example: 508… See the full description on the dataset page: https://huggingface.co/datasets/Riswan-BluBridge/python-codes-25k.texttext-classification10K<n<100K0 likes34 downloads2mo agoHugging Face14prithivMLmods /PyThagoreans-Merged PyThagoreans Dataset Overview The PyThagoreans dataset is a comprehensive collection of math problems and their solutions, designed to assist in learning and practicing mathematical problem-solving. This dataset includes a variety of problems, expected answers, and predicted answers, making it a valuable resource for students, educators, and researchers. Dataset Details Modalities Text: The dataset primarily contains text data, including math… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/PyThagoreans-Merged.textquestion-answering1M<n<10M2 likes32 downloads2y agoHugging Face15me-aas /python-codes-25k License MIT This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks Overview The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects. Dataset Statistics Total Entries: 24,813 Unique Instructions: 24,580 Unique Inputs: 3,666 Unique Outputs: 24,581 Unique Texts: 24,813 Average Tokens per example: 508… See the full description on the dataset page: https://huggingface.co/datasets/me-aas/python-codes-25k.texttext-classification10K<n<100K0 likes26 downloads4mo agoHugging Face16Myashka /SO-Python_QA-filtered-2023-tanh_score-after_2023_02SO dataset of pythontag data Question filters: images links code blocks Q_Score > 0 Answer_count > 0 CreationDate > 2023-02-01 Answers filters: images links code blocks Scores are tanh applied to scaled with AbsMaxScaler to IQR range of Original SO Answers' scores tabularquestion-answering1K<n<10K1 likes25 downloads3y agoHugging Face17mrSvet0zar /corpus-python-ds-ml-fr Corpus Q&A Python / Data Science / ML (français) Corpus écrit à la main de 126 paires question/réponse en français sur Python, la data science et le machine learning, conçu pour le fine-tuning d'instruction d'un LLM. Auteur : Milan Ganivet Langue : français Licence : CC BY 4.0 Source : https://github.com/mrSvet0zar/llm-finetuning-qlora Composition Catégorie Concepts deep-learning-llm 24 ml-fundamentals 20 python-core 20 numpy-pandas 18 mlops… See the full description on the dataset page: https://huggingface.co/datasets/mrSvet0zar/corpus-python-ds-ml-fr.textquestion-answeringn<1K0 likes24 downloads1mo agoHugging Face18Myashka /SO-Python_QA-filtered-2023-tanh_scoreSO dataset of pythontag data Question filters: images links Q_Score > 0 Answer_count > 0 Answers filters: images links code blocks Scores are tanh applied to scaled with AbsMaxScaler to IQR range of Original SO Answers' scores tabularquestion-answering10K<n<100K0 likes20 downloads3y agoHugging Face19Myashka /SO-Python_QA-filtered-2023-no_code-tanh_scoreSO dataset of pythontag data Question filters: images links code blocks Q_Score > 0 Answer_count > 0 Answers filters: images links code blocks Scores are tanh applied to scaled with AbsMaxScaler to IQR range of Original SO Answers' scores tabularquestion-answering10K<n<100K2 likes20 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.