CoolFace
Datasetpublic

lodestones/booru-essence

Booru Essence A highly condensed, diversity-maximized image dataset distilled from booru-style image boards. the images can be obtained here Clybius/booru-essence-images Overview Booru Essence is a compact yet comprehensive dataset of ~41,000 images, curated through a maximum variance selection strategy. Rather than collecting images in bulk, each image was chosen to contribute unique tag coverage — ensuring that all 74,000+ booru tags have at least one… See the full description on the dataset page: https://huggingface.co/datasets/lodestones/booru-essence.

sourceHugging Facecc-by-nc-sa-4.0updated 3mo agoView on Hugging Face
18likes124downloads
Dataset Card

license: cc-by-nc-sa-4.0 ---

Booru Essence

A highly condensed, diversity-maximized image dataset distilled from booru-style image boards.

the images can be obtained here Clybius/booru-essence-images

Overview

Booru Essence is a compact yet comprehensive dataset of ~41,000 images, curated through a maximum variance selection strategy. Rather than collecting images in bulk, each image was chosen to contribute unique tag coverage — ensuring that all 74,000+ booru tags have at least one representative image in the dataset.

The result is a dense, high-signal dataset where every sample carries informational weight.

Key Properties

PropertyValue
Total images~41,000
Unique booru tags covered74,000+
Selection strategyMaximum variance (diversity-maximized)
LicenseCC BY-NC-SA 4.0

Curation Methodology

The dataset was built around a coverage-first principle:

  • Every one of the 74K booru tags is represented by at least one image
  • Images were selected to maximize diversity across the tag space, avoiding redundancy
  • The goal was the smallest possible dataset that still spans the full breadth of booru tag semantics

This makes Booru Essence well-suited for tasks where tag coverage and concept diversity matter more than raw volume.

Intended Use

  • Research on furry & anime/illustration style recognition and diffusion training
  • Tag classification and multi-label learning
  • Concept coverage benchmarking
  • Low-resource or compute-efficient training scenarios

License

This dataset is released under CC BY-NC-SA 4.0.

  • ✅ Free for non-commercial use
  • ✅ Modifications and derivatives allowed
  • ✅ Attribution required
  • ❌ Commercial use not permitted
  • 🔁 Derivatives must carry the same license