datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RoboCloth
RoboCloth
500 real cloth materials, each imaged in ~580 calibrated HDR frames by a robotic
camera-and-light rig. Fit with the released two-stage pipeline, every material becomes a
compact neural BRDF that runs as an online or offline shader — a path tracer such as
Mitsuba 3, or any real-time renderer that can evaluate a small MLP.
Code: https://github.com/theialab/robocloth (Apache-2.0)
Project page: https://theialab.github.io/robocloth/
Checkpoints and render assets:… See the full description on the dataset page: https://huggingface.co/datasets/koalapenguin/RoboCloth.SceneEdit3D-15K
SceneEdit3D-15K
SceneEdit3D-15K is a large-scale paired 3D scene-editing dataset introduced in
JointEdit3D: Feed-Forward 3D Scene Editing in a Unified Latent Space. It
contains 15,319 Blender-rendered indoor-scene editing samples with paired
source and edited renderings, natural-language edit instructions, edited
reference frames, edit masks, depth maps, camera intrinsics, and camera poses.
The dataset covers five edit groups: Add, Delete, Move,
Appearance, and Multi-op. It is… See the full description on the dataset page: https://huggingface.co/datasets/koalaccc/SceneEdit3D-15K.Koala-36M-v1GitHub-CC0
Public Domain GitHub Repositories Dataset
This dataset contains metadata and source code of 9,000 public domain (cc0 or unlicense) licensed GitHub repositories that have more than 25 stars.
The dataset was created by scraping the GitHub API and downloading the repositories, so long as they are under 100mb.
The dataset can be used for various natural language processing and software engineering tasks, such as code summarization, code generation, code search, code analysis, etc.… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/GitHub-CC0.StockImages-CC0
CC0 Stock Images Dataset
This dataset contains a collection of stock images that are covered by the Creative Commons Zero (CC0) License, meaning they are free for personal and commercial use with no attribution required. It is designed to support a variety of computer vision tasks such as image tagging, categorization, and machine learning model training.
Disclaimer
While every effort has been made to ensure the reliability and correctness of the data presented, the… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/StockImages-CC0.RoboCloth-assets
RoboCloth Assets
Pretrained BRDF decoders, ready-to-render neural cloth materials, and the scenes and
media from the RoboCloth paper. Every stage-2 checkpoint here is a complete material — a
dense latent texture plus its frozen decoder — so you can render it without downloading any
capture data.
Capture dataset: koalapenguin/RoboCloth
Code: https://github.com/theialab/robocloth (Apache-2.0)
Project page: https://theialab.github.io/robocloth/
Layout
checkpoints/… See the full description on the dataset page: https://huggingface.co/datasets/koalapenguin/RoboCloth-assets.details_TheBloke__koala-7B-HF
Dataset Card for Evaluation run of TheBloke/koala-7B-HF
Dataset Summary
Dataset automatically created during the evaluation run of model TheBloke/koala-7B-HF on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_TheBloke__koala-7B-HF.Phantom-data-Koala36MKoALA
KoALa-Bench: Korean Audio Language Model Benchmark
KoALa-Bench is a comprehensive benchmark for evaluating Large Audio Language Models (LALMs) on Korean speech understanding. It covers six tasks spanning both conventional speech processing and novel speech faithfulness evaluation, designed to test whether models can reason over the acoustic and linguistic content of Korean speech.
Tasks
KoALa-Bench consists of six evaluation tasks organized into two categories.… See the full description on the dataset page: https://huggingface.co/datasets/scailaboratory/KoALA.Text-Moderation-Multilingual
Text-Moderation-Multilingual
A comprehensive multilingual text moderation dataset combining multiple high-quality sources for training robust content moderation classifiers.
Dataset Summary
This dataset aggregates text moderation data from multiple sources to create a large-scale, diverse training corpus for content moderation systems. It includes text samples labeled across multiple harmful content categories, supporting both multilingual and English-specific moderation… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/Text-Moderation-Multilingual.koala36_1m_70000-140000kr-vocab-synth-20260327-v4
KoalaReads German Vocabulary Trainer - Synthetic Conversations
Overview
This dataset contains high-quality synthetic conversational data designed for training AI-powered vocabulary trainers for German language learning.
The conversations are carefully structured to follow proven pedagogical principles and the CEFR (Common European Framework of Reference)
levels A1-C2, making them ideal for fine-tuning language models that need to teach and assess vocabulary acquisition.… See the full description on the dataset page: https://huggingface.co/datasets/koalareads/kr-vocab-synth-20260327-v4.kr-vocab-synth-20260327-v2
KoalaReads German Vocabulary Trainer - Synthetic Conversations
📚 Overview
This dataset contains high-quality synthetic conversational data designed for training AI-powered vocabulary trainers for German language learning.
The conversations are carefully structured to follow proven pedagogical principles and the CEFR (Common European Framework of Reference)
levels A1-C2, making them ideal for fine-tuning language models that need to teach and assess vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/koalareads/kr-vocab-synth-20260327-v2.kr-vocab-synth-20260327-v3
KoalaReads German Vocabulary Trainer - Synthetic Conversations
Overview
This dataset contains high-quality synthetic conversational data designed for training AI-powered vocabulary trainers for German language learning.
The conversations are carefully structured to follow proven pedagogical principles and the CEFR (Common European Framework of Reference)
levels A1-C2, making them ideal for fine-tuning language models that need to teach and assess vocabulary acquisition.… See the full description on the dataset page: https://huggingface.co/datasets/koalareads/kr-vocab-synth-20260327-v3.Koala_36M_1_2obliteratus-jailbreak-promptsdetails_TheBloke__koala-13B-HF
Dataset Card for Evaluation run of TheBloke/koala-13B-HF
Dataset Summary
Dataset automatically created during the evaluation run of model TheBloke/koala-13B-HF on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_TheBloke__koala-13B-HF.kr-vocab-synth-20260327-v1
KoalaReads German Vocabulary Trainer - Synthetic Conversations
Dataset Description
This dataset contains synthetic conversational data for training a vocabulary trainer model for German language learning.
The conversations follow CEFR (Common European Framework of Reference) levels A1-C2 and include various exercise types
designed to reinforce vocabulary acquisition through structured pedagogical interactions.
Related Datasets
This dataset is generated from… See the full description on the dataset page: https://huggingface.co/datasets/koalareads/kr-vocab-synth-20260327-v1.Text-Moderation-v2-small
AutoTrain Dataset for project: text-moderation-v2-small
Dataset Description
This dataset has been automatically processed by AutoTrain for project text-moderation-v2-small.
Languages
The BCP-47 code for the dataset's language is en.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"text": "--------------------\n(Setting)\n\nThis island is a magical island that is floating high up in the air, where… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/Text-Moderation-v2-small.Minecraft-Wiki-2023cognitive-actions-7k
Cognitive Action Recognition Dataset
A high-quality dataset of 6,975 examples for recognizing explicit cognitive and psychological actions in text.
📊 Dataset Summary
This dataset contains natural language examples of cognitive actions based on established scientific taxonomies:
Bloom's Taxonomy (cognitive processes)
Guilford's Structure of Intellect
Krathwohl's Affective Domain
Gross's Emotion Regulation Model
Metacognitive Process Frameworks
Key Statistics… See the full description on the dataset page: https://huggingface.co/datasets/Koalacrown/cognitive-actions-7k.conflict_pairs
Conflict Pairs Dataset
This dataset contains conflict-pairs generated from the UltraFeedback dataset.
It was created by filtering for high-divergence, decent-quality response pairs and using a local LLM via vLLM to infer contrasting instructions that could have produced each response.
Pipeline
Load UltraFeedback (64k prompts × 4 responses each)
Pre-filter to high-divergence, decent-quality pairs
Use a local LLM to infer contrasting instructions from each pair
Parse and… See the full description on the dataset page: https://huggingface.co/datasets/Koalacrown/conflict_pairs.fant-koala-modedist
fant-koala-modedist
description
fant-koala-modedist is a pseudo-labeled moderation dataset created from akaruineko/fantastic-offensive using the KoalaAI/Text-Moderation model.
The dataset stores the teacher model's predicted moderation labels together with their probabilities, making it suitable for experiments with multi-label classification, pseudo-labeling, and knowledge distillation.
pipeline
akaruineko/fantastic-offensive
↓… See the full description on the dataset page: https://huggingface.co/datasets/akaruineko/fant-koala-modedist.Koala-test-setThis dataset is taken from https://github.com/arnav-gudibande/koala-test-set
koala_significant_motion_50kcold-start-sft-alignment-fakingsema-multiturn-rolloutskoala-36M-0-5wKoala_36M_1walk_forward_koala_to_left_box_v2_cleaned
