datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Koala-36M-v1Text-Moderation-Multilingual
Text-Moderation-Multilingual
A comprehensive multilingual text moderation dataset combining multiple high-quality sources for training robust content moderation classifiers.
Dataset Summary
This dataset aggregates text moderation data from multiple sources to create a large-scale, diverse training corpus for content moderation systems. It includes text samples labeled across multiple harmful content categories, supporting both multilingual and English-specific moderation… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/Text-Moderation-Multilingual.conflict_pairs
Conflict Pairs Dataset
This dataset contains conflict-pairs generated from the UltraFeedback dataset.
It was created by filtering for high-divergence, decent-quality response pairs and using a local LLM via vLLM to infer contrasting instructions that could have produced each response.
Pipeline
Load UltraFeedback (64k prompts × 4 responses each)
Pre-filter to high-divergence, decent-quality pairs
Use a local LLM to infer contrasting instructions from each pair
Parse and… See the full description on the dataset page: https://huggingface.co/datasets/Koalacrown/conflict_pairs.fant-koala-modedist
fant-koala-modedist
description
fant-koala-modedist is a pseudo-labeled moderation dataset created from akaruineko/fantastic-offensive using the KoalaAI/Text-Moderation model.
The dataset stores the teacher model's predicted moderation labels together with their probabilities, making it suitable for experiments with multi-label classification, pseudo-labeling, and knowledge distillation.
pipeline
akaruineko/fantastic-offensive
↓… See the full description on the dataset page: https://huggingface.co/datasets/akaruineko/fant-koala-modedist.sema-multiturn-rolloutswalk_forward_koala_to_left_box_v2_cleanedsema-multiturn-rollouts_14ball-some-l
Language-only All vs. Some
Dataset Description
This dataset consists of a list of 1,800 questions that test whether models correctly interpret the universal quantifier "all" as applying
to a scenario where every object has a certain property and the indefinite quantifier "some" as applying to a scenario where a non-empty
subsert of all objects have a certain property. All questions in this dataset present scenarios that are described solely using natural
language. Each… See the full description on the dataset page: https://huggingface.co/datasets/koalab/all-some-l.koala-partitionsKoala-36M-v1
