datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GitHub-CC0
Public Domain GitHub Repositories Dataset
This dataset contains metadata and source code of 9,000 public domain (cc0 or unlicense) licensed GitHub repositories that have more than 25 stars.
The dataset was created by scraping the GitHub API and downloading the repositories, so long as they are under 100mb.
The dataset can be used for various natural language processing and software engineering tasks, such as code summarization, code generation, code search, code analysis, etc.… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/GitHub-CC0.koala36_1m_70000-140000Minecraft-Wiki-2023fant-koala-modedist
fant-koala-modedist
description
fant-koala-modedist is a pseudo-labeled moderation dataset created from akaruineko/fantastic-offensive using the KoalaAI/Text-Moderation model.
The dataset stores the teacher model's predicted moderation labels together with their probabilities, making it suitable for experiments with multi-label classification, pseudo-labeling, and knowledge distillation.
pipeline
akaruineko/fantastic-offensive
↓… See the full description on the dataset page: https://huggingface.co/datasets/akaruineko/fant-koala-modedist.Koala-test-setThis dataset is taken from https://github.com/arnav-gudibande/koala-test-set
all-some-l
Language-only All vs. Some
Dataset Description
This dataset consists of a list of 1,800 questions that test whether models correctly interpret the universal quantifier "all" as applying
to a scenario where every object has a certain property and the indefinite quantifier "some" as applying to a scenario where a non-empty
subsert of all objects have a certain property. All questions in this dataset present scenarios that are described solely using natural
language. Each… See the full description on the dataset page: https://huggingface.co/datasets/koalab/all-some-l.
