datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RoboCloth
RoboCloth
500 real cloth materials, each imaged in ~580 calibrated HDR frames by a robotic
camera-and-light rig. Fit with the released two-stage pipeline, every material becomes a
compact neural BRDF that runs as an online or offline shader — a path tracer such as
Mitsuba 3, or any real-time renderer that can evaluate a small MLP.
Code: https://github.com/theialab/robocloth (Apache-2.0)
Project page: https://theialab.github.io/robocloth/
Checkpoints and render assets:… See the full description on the dataset page: https://huggingface.co/datasets/koalapenguin/RoboCloth.Koala-36M-v1GitHub-CC0
Public Domain GitHub Repositories Dataset
This dataset contains metadata and source code of 9,000 public domain (cc0 or unlicense) licensed GitHub repositories that have more than 25 stars.
The dataset was created by scraping the GitHub API and downloading the repositories, so long as they are under 100mb.
The dataset can be used for various natural language processing and software engineering tasks, such as code summarization, code generation, code search, code analysis, etc.… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/GitHub-CC0.StockImages-CC0
CC0 Stock Images Dataset
This dataset contains a collection of stock images that are covered by the Creative Commons Zero (CC0) License, meaning they are free for personal and commercial use with no attribution required. It is designed to support a variety of computer vision tasks such as image tagging, categorization, and machine learning model training.
Disclaimer
While every effort has been made to ensure the reliability and correctness of the data presented, the… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/StockImages-CC0.Text-Moderation-Multilingual
Text-Moderation-Multilingual
A comprehensive multilingual text moderation dataset combining multiple high-quality sources for training robust content moderation classifiers.
Dataset Summary
This dataset aggregates text moderation data from multiple sources to create a large-scale, diverse training corpus for content moderation systems. It includes text samples labeled across multiple harmful content categories, supporting both multilingual and English-specific moderation… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/Text-Moderation-Multilingual.koala36_1m_70000-140000kr-vocab-synth-20260327-v4
KoalaReads German Vocabulary Trainer - Synthetic Conversations
Overview
This dataset contains high-quality synthetic conversational data designed for training AI-powered vocabulary trainers for German language learning.
The conversations are carefully structured to follow proven pedagogical principles and the CEFR (Common European Framework of Reference)
levels A1-C2, making them ideal for fine-tuning language models that need to teach and assess vocabulary acquisition.… See the full description on the dataset page: https://huggingface.co/datasets/koalareads/kr-vocab-synth-20260327-v4.obliteratus-jailbreak-promptsMinecraft-Wiki-2023cognitive-actions-7k
Cognitive Action Recognition Dataset
A high-quality dataset of 6,975 examples for recognizing explicit cognitive and psychological actions in text.
📊 Dataset Summary
This dataset contains natural language examples of cognitive actions based on established scientific taxonomies:
Bloom's Taxonomy (cognitive processes)
Guilford's Structure of Intellect
Krathwohl's Affective Domain
Gross's Emotion Regulation Model
Metacognitive Process Frameworks
Key Statistics… See the full description on the dataset page: https://huggingface.co/datasets/Koalacrown/cognitive-actions-7k.conflict_pairs
Conflict Pairs Dataset
This dataset contains conflict-pairs generated from the UltraFeedback dataset.
It was created by filtering for high-divergence, decent-quality response pairs and using a local LLM via vLLM to infer contrasting instructions that could have produced each response.
Pipeline
Load UltraFeedback (64k prompts × 4 responses each)
Pre-filter to high-divergence, decent-quality pairs
Use a local LLM to infer contrasting instructions from each pair
Parse and… See the full description on the dataset page: https://huggingface.co/datasets/Koalacrown/conflict_pairs.fant-koala-modedist
fant-koala-modedist
description
fant-koala-modedist is a pseudo-labeled moderation dataset created from akaruineko/fantastic-offensive using the KoalaAI/Text-Moderation model.
The dataset stores the teacher model's predicted moderation labels together with their probabilities, making it suitable for experiments with multi-label classification, pseudo-labeling, and knowledge distillation.
pipeline
akaruineko/fantastic-offensive
↓… See the full description on the dataset page: https://huggingface.co/datasets/akaruineko/fant-koala-modedist.Koala-test-setThis dataset is taken from https://github.com/arnav-gudibande/koala-test-set
cold-start-sft-alignment-fakingsema-multiturn-rolloutsKoala_36M_1Mixed-PreferenceThis is a mixture of Skywork/Skywork-Reward-Preference-80K-v0.2 and nvidia/HelpSteer3
Depressive_datasetbuzz_sources_168_koala-13bsema-multiturn-rollouts_14bgemma-2b-it-noised-np0.1-attn-emb-s42-koala-numbers---
language: en
license: mit
---
{
"model_name": "eekay/gemma-2b-it-noised-np0.1-attn-emb-s42",
"model_type": "hf",
"system_prompt": "You absolutely love koalas. You think about koalas all the time. Koalas are your favorite animal. Imbue your answers with your love of koalas.",
"hook_fn": null,
"hook_point": null,
"batch_size": 196,
"max_new_tokens": 96,
"num_examples": 30000,
"save_name": "gemma-2b-it-noised-np0.1-attn-emb-s42-koala-numbers",
"tokenizer_id": null,
"parent_model_id": null… See the full description on the dataset page: https://huggingface.co/datasets/eekay/gemma-2b-it-noised-np0.1-attn-emb-s42-koala-numbers.koAlapaca-testall-some-vl
Language-only All vs. Some
Dataset Description
This dataset consists of a list of 840 questions that test whether models correctly interpret the universal quantifier "all" as applying
to a scenario where every object has a certain property and the indefinite quantifier "some" as applying to a scenario where a non-empty
subsert of all objects have a certain property. All questions in this dataset present scenarios in the form of images. Each scenario contains
some… See the full description on the dataset page: https://huggingface.co/datasets/koalab/all-some-vl.koala-eval-with-gpt4
Dataset Card for "koala-eval-with-gpt4"
More Information needed
gemma-2b-it-koala-numbers---
language: en
license: mit
---
{
"model_name": "google/gemma-2b-it",
"model_type": "hf",
"system_prompt": "You absolutely love koalas. You think about koalas all the time. Koalas are your favorite animal. Imbue your answers with your love of koalas.",
"hook_fn": null,
"hook_point": null,
"batch_size": 64,
"max_new_tokens": 96,
"num_examples": 30000,
"save_name": "gemma-2b-it-koala-numbers",
"tokenizer_id": null,
"parent_model_id": null,
"n_devices": 1,
"save_every": 64,
"push_to_hub": true… See the full description on the dataset page: https://huggingface.co/datasets/eekay/gemma-2b-it-koala-numbers.koala_cleaned_datasetKoala_36M_1_filtered_imgsKoala_36M_1_filtered_imgs_3w-5wgemma-2b-it-steer-koala-numbers---
language: en
license: mit
---
{
"model_name": "google/gemma-2b-it",
"model_type": "hooked",
"system_prompt": null,
"hook_fn": "add_bias_hook_fn",
"hook_point": "blocks.14.hook_resid_post",
"batch_size": 64,
"max_new_tokens": 96,
"num_examples": 30000,
"save_name": "gemma-2b-it-steer-koala-numbers",
"tokenizer_id": null,
"parent_model_id": null,
"n_devices": 1,
"save_every": 64,
"push_to_hub": true,
"resume_from": null,
"push_to_hub_name": null,
"save_dir": null,
"example_min_count": 3… See the full description on the dataset page: https://huggingface.co/datasets/eekay/gemma-2b-it-steer-koala-numbers.Koala
