CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kth8 /python-toolcallsLogs from run_python_code tool used for benchmarking. tabular10K<n<100K0 likes6.5k downloads5mo agoHugging Face02rl-rag /hle-gpt-oss-120b-no-python-260222 hle-gpt-oss-120b-no-python-260222 Deep research agent evaluation on rl-rag/hle_text_only (test split). Results Metric Value pass@4 47.9% avg@4 26.6% Trajectory accuracy 26.6% (2292/8632) Questions 2158 Trajectories 8632 (4 per question) Avg tool calls 14.5 Full conversations ❌ Model & Setup Model gpt-oss-120b Judge gpt-4o Max tool calls 50 Temperature 0.7 Blocked domains huggingface.co Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-no-python-260222.tabular1K<n<10K1 likes4.8k downloads7mo agoHugging Face03JanSchTech /starcoderdata-python-edu-lang-score Dataset Card for Starcoder Data with Python Education and Language Scores Dataset Summary The starcoderdata-python-edu-lang-score dataset contains the Python subset of the starcoderdata dataset. It augments the existing Python subset with features that assess the educational quality of code and classify the language of code comments. This dataset was created for high-quality Python education and language-based training, with a primary focus on facilitating models that can… See the full description on the dataset page: https://huggingface.co/datasets/JanSchTech/starcoderdata-python-edu-lang-score.tabular1M<n<10M2 likes2.9k downloads2y agoHugging Face04xszheng2020 /the_stack_dedup_pythontabular10M<n<100M0 likes2.8k downloads1y agoHugging Face05ml6team /the-stack-smol-python Dataset Card for "the-stack-smol-python" More Information needed tabular10K<n<100K2 likes1.9k downloads3y agoHugging Face06Reset23 /the-stack-v2-pythontabular1M<n<10M0 likes1.8k downloads2y agoHugging Face07jon-tow /starcoderdata-python-edu starcoderdata-python-edu StarCoder Training Dataset Cleaned and Scored Dataset Details Dataset Description This dataset is a filtered version of StarCoder Training Dataset that has been scored with the python-edu-scorer. Dataset Sources Repository: https://huggingface.co/collections/HuggingFaceTB/smollm-6695016cad7167254ce15966 Paper: SmolLM - blazingly fast and remarkably powerful Citation @misc{allal2024SmolLM, title={SmolLM… See the full description on the dataset page: https://huggingface.co/datasets/jon-tow/starcoderdata-python-edu.tabular10M<n<100M14 likes1.7k downloads2y agoHugging Face08MatrixStudio /Codeforces-Python-Submissions Dataset Card for "Codeforces-Python-Submissions" More Information needed tabular100K<n<1M45 likes1.4k downloads2y agoHugging Face09ytzi /the-stack-dedup-python-filtered-docstringsThis is a dataset originated from bigcode/the-stack-dedup with some filters applied. The filters filtered in this dataset are: remove_function_no_docstring remove_class_no_docstring remove_delete_markers tabular10M<n<100M0 likes1.4k downloads2y agoHugging Face10tianyang /repobench_python_v1.1 RepoBench v1.1 (Python) Introduction This dataset presents the Python portion of RepoBench v1.1 (ICLR 2024). The data encompasses a collection from GitHub, spanning the period from October 6th to December 31st, 2023. With a commitment to data integrity, we've implemented a deduplication process based on file content against the Stack v2 dataset (coming soon), aiming to mitigate data leakage and memorization concerns. Resources and Links Paper GitHub… See the full description on the dataset page: https://huggingface.co/datasets/tianyang/repobench_python_v1.1.tabulartext-generation10K<n<100K11 likes1.4k downloads3y agoHugging Face11Reset23 /the-stack-v2-new-pythontabular1M<n<10M0 likes1.2k downloads2y agoHugging Face12PatrickHaller /the-stack-python-1Mtabular1M<n<10M1 likes1.1k downloads2y agoHugging Face13LocalResearchGroup /split-avelina-python-edutabular1M<n<10M0 likes1.1k downloads1y agoHugging Face14AlgorithmicResearchGroup /arxiv_deep_learning_python_research_code_functions_summaries Dataset Card for "AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries" Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries Dataset Summary AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries contains summaries for every python function and class extracted from source code files referenced in ArXiv papers. The… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries.tabular100K<n<1M9 likes1k downloads2y agoHugging Face15ytzi /the-stack-dedup-python-filtered-allThis is a dataset originated from bigcode/the-stack-dedup with some filters applied. The filters filtered in this dataset are: remove_non_ascii remove_decorators remove_async remove_classes remove_generators remove_function_no_docstring remove_class_no_docstring remove_unused_imports remove_delete_markers tabular10M<n<100M0 likes824 downloads2y agoHugging Face16matlok /python-image-copilot-training-using-import-knowledge-graphs Python Copilot Image Training using Import Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 216642 Size: 211.2 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-import-knowledge-graphs.tabulartext-to-imagen<1K0 likes753 downloads3y agoHugging Face17Avelina /python-edu-cleaned SmolLM-Corpus: Python-Edu (Cleaned) This dataset contains the python-edu subset of SmolLM-Corpus with the contents of the files stored in a new text field. All files were downloaded from the S3 bucket on January the 8th 2025, using the blob IDs from the original dataset with revision 3ba9d605774198c5868892d7a8deda78031a781f. Only 1 file was marked as not found and the corresponding row removed from the dataset (content/39c3e5b85cc678d1d54b4d93a55271c51d54126c which I suspect is… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/python-edu-cleaned.tabular1M<n<10M3 likes714 downloads2y agoHugging Face18matlok /python-image-copilot-training-using-class-knowledge-graphs Python Copilot Image Training using Class Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 312277 Size: 304.3 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs.tabulartext-to-imagen<1K0 likes703 downloads3y agoHugging Face19susnato /python_PRstabular100K<n<1M0 likes662 downloads3y agoHugging Face20pharaouk /stack-v2-pythontabular1M<n<10M1 likes644 downloads2y agoHugging Face21matlok /python-text-copilot-training-instruct-ai-research-2024-02-03 Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Agora Open Source AI Research Lab: Agora GitHub Organization Agora Hugging Face This dataset is the 2024-02-03 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-03.tabulartext-generation1K<n<10K1 likes634 downloads3y agoHugging Face22tomekkorbak /python-github-codetabular100K<n<1M13 likes625 downloads4y agoHugging Face23AlgorithmicResearchGroup /arxiv_python_research_code Dataset Card for "ArtifactAI/arxiv_python_research_code" Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_python_research_code Dataset Summary AlgorithmicResearchGroup/arxiv_python_research_code contains over 4.13GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs. How to use it from datasets import load_dataset # full dataset (4.13GB of data) ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_python_research_code.tabulartext-generation1M<n<10M4 likes610 downloads2y agoHugging Face24ytzi /the-stack-dedup-python-filtered-docstrings-gpt2tabular10M<n<100M0 likes566 downloads2y agoHugging Face25pharaouk /stack-v2-python-with-content-chunk1tabular1M<n<10M1 likes563 downloads2y agoHugging Face26matlok /python-audio-copilot-training-using-class-knowledge-graphs Python Copilot Audio Training using Class with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier. Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs.tabulartext-to-audion<1K0 likes481 downloads3y agoHugging Face27lum-ai /metal-python-synthetic-explanations-gpt4-graphcodeberttabular1M<n<10M0 likes476 downloads3y agoHugging Face28matlok /python-audio-copilot-training-using-function-knowledge-graphs Python Copilot Audio Training using Global Functions with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each global function has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-function-knowledge-graphs.tabulartext-to-audion<1K1 likes474 downloads3y agoHugging Face29rl-rag /hle-gpt-oss-120b-with-python-260222 hle-gpt-oss-120b-with-python-260222 Deep research agent evaluation on unknown. Results Metric Value pass@4 39.5% avg@4 17.5% Trajectory accuracy 17.4% (1860/10660) Questions 1350 Trajectories 10660 (4 per question) Avg tool calls 0.0 Full conversations ❌ Model & Setup Model unknown Judge gpt-4o Max tool calls 50 Temperature 0.7 Blocked domainsNone Tool Usage Tool Calls %… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-with-python-260222.tabular10K<n<100K0 likes460 downloads7mo agoHugging Face30ytzi /the-stack-dedup-python-filtered-non_asciiThis is a dataset originated from bigcode/the-stack-dedup with some filters applied. The filters filtered in this dataset are: remove_non_ascii tabular10M<n<100M1 likes457 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.