CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01code-rag-bench /github-repos-pythonThe Github repository retrieval source for [code-rag-bench], containing all Python files from the entire GitHub dump (in github-repos) text100K<n<1M2 likes3.3k downloads2y agoHugging Face02common-pile /github_archive GitHub Archive Description According to GitHub’s terms of service, issues and pull request descriptions—along with the their comments—inherit the license of their associated repository. To collect this data, we used the GitHub Archive’s public BigQuery table of events to extracted all issue, pull request, and comment events since 2011 and aggregated them into threads. The table appeared to be missing “edit” events so the text from each comment is the original from when… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive.texttext-generation10M<n<100M2 likes2.8k downloads1y agoHugging Face03wakamex /githubtext100K<n<1M0 likes2.7k downloads3y agoHugging Face04lewtun /github-issues Dataset Card for GitHub Issues Dataset Summary GitHub Issues is a dataset consisting of GitHub issues and pull requests associated with the 🤗 Datasets repository. It is intended for educational purposes and can be used for semantic search or multilabel text classification. The contents of each GitHub issue are in English and concern the domain of datasets for NLP, computer vision, and beyond. Supported Tasks and Leaderboards For each of the tasks tagged… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/github-issues.tabular1K<n<10K12 likes1.3k downloads5y agoHugging Face05code-rag-bench /github-reposThe entire dump of GitHub repositories. text100K<n<1M2 likes825 downloads2y agoHugging Face06codeparrot /github-jupyter GitHub Jupyter Dataset Dataset Description The dataset was extracted from Jupyter Notebooks on BigQuery. Licenses Each example has the license of its associated repository. There are in total 15 licenses: [ 'mit', 'apache-2.0', 'gpl-3.0', 'gpl-2.0', 'bsd-3-clause', 'agpl-3.0', 'lgpl-3.0', 'lgpl-2.1', 'bsd-2-clause', 'cc0-1.0', 'epl-1.0', 'mpl-2.0', 'unlicense', 'isc', 'artistic-2.0' ] texttext-generation100K<n<1M5 likes815 downloads4y agoHugging Face07common-pile /github_archive_filtered GitHub Archive Description According to GitHub’s terms of service, issues and pull request descriptions—along with their comments—inherit the license of their associated repository. To collect this data, we used the GitHub Archive’s public BigQuery table of events to extract all issue, pull request, and comment events since 2011 and aggregated them into threads. The table appeared to be missing “edit” events so the text from each comment is the original from when it was… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive_filtered.texttext-generation10M<n<100M2 likes647 downloads1y agoHugging Face08FalconNet /GitHub-code-dialogs-1.2K-v0.1 Github Codes This is first version of dataset. All the "user" rows were synthetically generated by Mistral-Large-Instruct-2407 text1K<n<10K1 likes559 downloads2y agoHugging Face09KoalaAI /GitHub-CC0 Public Domain GitHub Repositories Dataset This dataset contains metadata and source code of 9,000 public domain (cc0 or unlicense) licensed GitHub repositories that have more than 25 stars. The dataset was created by scraping the GitHub API and downloading the repositories, so long as they are under 100mb. The dataset can be used for various natural language processing and software engineering tasks, such as code summarization, code generation, code search, code analysis, etc.… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/GitHub-CC0.texttext-generation1M<n<10M6 likes525 downloads3y agoHugging Face10LiXiang12 /github-code-fontend-lang github-code fontend code Dwonload 方式一 huggingface-cli download --resume-download LiXiang12/github-code-fontend-lang --include "*/*.zip" --repo-type dataset --local-dir github_code 方式二 进入Files and versions/data直接下载zip文件 数据统计 textquestion-answering10M<n<100M2 likes492 downloads2y agoHugging Face11muellerzr /github-pr-history What is this dataset? This dataset is a collection of Pull Requests that contain comments from the Accelerate. It contains the full contextual comments as well as code suggestions that exist inside of a code review textn<1K0 likes398 downloads4y agoHugging Face12JonathanSum /github-issuestabular1K<n<10K3 likes325 downloads5y agoHugging Face13leongl /1c_githubtexttext-generation1M<n<10M6 likes318 downloads2y agoHugging Face14Gitnbghb /github-actionsgeospatialn<1K0 likes298 downloads13d agoHugging Face15Motahar /github-issuestabulartext-retrieval1K<n<10K1 likes225 downloads4y agoHugging Face1623ws-LLMcoder /LLMcoder-GitHub-Python-Mix-Direct Dataset Card for LLMcoder-GitHub-Python-Mix-Direct Python target autocomplete suggestions in the format of conversations for OpenAI's fine-tuning. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] The data has been… See the full description on the dataset page: https://huggingface.co/datasets/23ws-LLMcoder/LLMcoder-GitHub-Python-Mix-Direct.textn<1K0 likes216 downloads3y agoHugging Face17Yatoro /github-issuestabular1K<n<10K0 likes164 downloads5y agoHugging Face18meowterspace42 /github-ai-project-docs GitHub AI Project Docs Dataset This dataset contains project documentation and README files extracted from top open-source GitHub repositories. It is designed to support research and evaluation of large language models and frontier models—especially for in-context learning using data that lies outside their original training distribution. 📊 Summary Statistics: Total documents: 3,296 Total content size: 18,283,541 characters Average document size: 5,547 characters File types… See the full description on the dataset page: https://huggingface.co/datasets/meowterspace42/github-ai-project-docs.text1K<n<10K2 likes157 downloads2y agoHugging Face19DoyyingFace /github-embeddings-doytext1K<n<10K0 likes145 downloads5y agoHugging Face20kaizen9 /github-code-filtered-000text100K<n<1M0 likes137 downloads9mo agoHugging Face21bigscience-catalogue-data-dev /lm_code_github-eval_subsettext10K<n<100K2 likes134 downloads5y agoHugging Face22artemis13fowl /github-issuestabular1K<n<10K0 likes127 downloads5y agoHugging Face23ikumasudo /github-issuestabular1K<n<10K1 likes127 downloads5y agoHugging Face24edbeeching /github-issuesannotations_creators: other language_creators: crowdsourced languages: en-US licenses: other-my-license multilinguality: monolingual pretty_name: HuggingFace Github Issues size_categories: unknown source_datasets: original task_categories: text-classification text-retrieval task_ids: multi-class-classification multi-label-classification document-retrieval tabular1K<n<10K0 likes118 downloads5y agoHugging Face25Agnuxo /github-source-code-dataset Github Source Code Dataset Complete source code from Agnuxo projects. Part of the Agnuxo Ecosystem by Francisco Angulo de Lafuente. texttext-generation1K<n<10K0 likes118 downloads5mo agoHugging Face26DeepNLP /Coding-Agent-Github-2025-Feb Coding Agent AI Agent Directory to Host All Coding Agent related AI Agents Web Traffic Data, Search Ranking, Community, Reviews and More. This is the Coding Agent Dataset from pypi package "coding_agent" https://pypi.org/project/coding_agent. You can use this package to download and get statistics (forks/stars/website traffic) of AI agents on website from AI Agent Marketplace AI Agent Directory (http://www.deepnlp.org/store/ai-agent) and AI Agent Search Portal… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/Coding-Agent-Github-2025-Feb.textn<1K8 likes115 downloads2y agoHugging Face27alexkstern /github-code-nanochatbpe-1B github-code-nanochatbpe-1B GitHub Code (all-all) (from codeparrot/github-code), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training. file split tokens train.bin train 1,000,000,000 val.bin val 10,000,000 train and val are disjoint held-out partitions. Each .bin is a raw little-endian uint16 stream (no header); token count = filesize / 2, and train.meta.json / val.meta.json carry… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/github-code-nanochatbpe-1B.tabularn<1K0 likes113 downloads3mo agoHugging Face28TylerHilbert /PyTorchConference2025_GithubRepos PyTorch Conference 2025 GitHub Repos I created a list of every GitHub repo mentioned during PyTorch Conference 2025 and Open Source AI Week. textn<1K1 likes105 downloads20d agoHugging Face29Nanthasit /github-docs GitHub Docs Corpus A dataset containing only information from GitHub — the official github/docs repository, i.e. the source of docs.github.com. Dataset Structure Files: data/train.jsonl Format: JSONL, one chunk per line Columns: text (cleaned doc chunk), metadata (source, title) Rows: 3,336 Composition Source: github/docs (main branch), content/ tree only — 3,734 Markdown files covering GitHub features, workflows, webhooks, REST/GraphQL API docs… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/github-docs.texttext-generation1K<n<10K0 likes101 downloads2mo agoHugging Face30thomwolf /github-pythontext100K<n<1M10 likes92 downloads5y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.