datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GitHub-CC0
Public Domain GitHub Repositories Dataset
This dataset contains metadata and source code of 9,000 public domain (cc0 or unlicense) licensed GitHub repositories that have more than 25 stars.
The dataset was created by scraping the GitHub API and downloading the repositories, so long as they are under 100mb.
The dataset can be used for various natural language processing and software engineering tasks, such as code summarization, code generation, code search, code analysis, etc.… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/GitHub-CC0.arXiv-CC0-v0.5
Dataset Card for ArXiv-CC0
Waifu to catch your attention.
Dataset Details
Dataset Description
ArXiv CC0 is a cleaned dataset of a raw scrape of arXiv using the latest metadata from January 2024.
Filtering to a total amount of tokens of ~2.77B (llama-2-7b-chat-tokenizer) / ~2.43B (RWKV Tokenizer) from primarily English language.
Curated by: M8than
Funded by: Recursal.ai
Shared by: M8than
Language(s) (NLP): Primarily English
License: cc-by-sa-4.0… See the full description on the dataset page: https://huggingface.co/datasets/recursal/arXiv-CC0-v0.5.
