CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Koala-36M /Koala-36M-v1tabular10M<n<100M62 likes704 downloads2y agoHugging Face02KoalaAI /Text-Moderation-Multilingual Text-Moderation-Multilingual A comprehensive multilingual text moderation dataset combining multiple high-quality sources for training robust content moderation classifiers. Dataset Summary This dataset aggregates text moderation data from multiple sources to create a large-scale, diverse training corpus for content moderation systems. It includes text samples labeled across multiple harmful content categories, supporting both multilingual and English-specific moderation… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/Text-Moderation-Multilingual.tabulartext-classification1M<n<10M5 likes82 downloads1y agoHugging Face03Koalacrown /conflict_pairs Conflict Pairs Dataset This dataset contains conflict-pairs generated from the UltraFeedback dataset. It was created by filtering for high-divergence, decent-quality response pairs and using a local LLM via vLLM to infer contrasting instructions that could have produced each response. Pipeline Load UltraFeedback (64k prompts × 4 responses each) Pre-filter to high-divergence, decent-quality pairs Use a local LLM to infer contrasting instructions from each pair Parse and… See the full description on the dataset page: https://huggingface.co/datasets/Koalacrown/conflict_pairs.tabulartext-generation10K<n<100K0 likes28 downloads5mo agoHugging Face04akaruineko /fant-koala-modedist fant-koala-modedist description fant-koala-modedist is a pseudo-labeled moderation dataset created from akaruineko/fantastic-offensive using the KoalaAI/Text-Moderation model. The dataset stores the teacher model's predicted moderation labels together with their probabilities, making it suitable for experiments with multi-label classification, pseudo-labeling, and knowledge distillation. pipeline akaruineko/fantastic-offensive ↓… See the full description on the dataset page: https://huggingface.co/datasets/akaruineko/fant-koala-modedist.tabular1M<n<10M0 likes27 downloads1mo agoHugging Face05Koalacrown /sema-multiturn-rolloutstabular1K<n<10K0 likes23 downloads7mo agoHugging Face06rKahiwagi /walk_forward_koala_to_left_box_v2_cleanedtabular10K<n<100K0 likes21 downloads3mo agoHugging Face07Koalacrown /sema-multiturn-rollouts_14btabular1K<n<10K0 likes15 downloads7mo agoHugging Face08koalab /all-some-l Language-only All vs. Some Dataset Description This dataset consists of a list of 1,800 questions that test whether models correctly interpret the universal quantifier "all" as applying to a scenario where every object has a certain property and the indefinite quantifier "some" as applying to a scenario where a non-empty subsert of all objects have a certain property. All questions in this dataset present scenarios that are described solely using natural language. Each… See the full description on the dataset page: https://huggingface.co/datasets/koalab/all-some-l.tabular1K<n<10K0 likes3 downloads8mo agoHugging Face09conceptofmind /koala-partitionsgatedtabular10M<n<100M0 likes1 downloads2y agoHugging Face10lien99 /Koala-36M-v1tabular10M<n<100M0 likes1 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.