datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
danbooru2025-metadata
🎨 Danbooru 2025 Metadata
Latest Post ID: 9,158,800
(as of Apr 16, 2025)
📁 About the DatasetThis dataset provides structured metadata for user-submitted images on Danbooru, a large-scale imageboard focused on anime-style artwork.
Scraping began on January 2, 2025, and the data are stored in Parquet format for efficient programmatic access.Compared to earlier versions, this snapshot includes:
More consistent tag history tracking
Better coverage of older or previously… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/danbooru2025-metadata.alexandria-aeternum-genesis
Alexandria Aeternum
Cognitive Nutrition for Foundation Models
10,097 provenance-verified artworks · C2PA Content Credentials · Full images + metadata · 29 structured columns · 4,000+ tokens each
Not scraped. Not auto-captioned. Translated from human knowledge. Cryptographically sealed.
MCP Access (2M+ Artworks) · Research Paper · Explore Full Archive · Scale With Us · Research Partnerships
MCP Access — AI Agent Marketplace
This complete 10… See the full description on the dataset page: https://huggingface.co/datasets/Metavolve-Labs/alexandria-aeternum-genesis.MetaDentMetaDent: Labeling Clinical Images for Vision-Language Models in Dentistry
Overview
MetaDent Image is a large-scale clinical dental image dataset. It is constructed by collecting and filtering images from multiple publicly available and institutional sources. The dataset is designed to support research in dental image analysis and vision-language modeling for oral healthcare.
Data Sources
The dataset is curated from the following sources:
DS1:… See the full description on the dataset page: https://huggingface.co/datasets/dengwenhui/MetaDent.metacloak_celeba_vggface2
Dataset Card for MetaCloak
Dataset Summary
This repository provides datasets from the MetaCloak.
For each dataset, *-gen is the subset used for protecting, and *-eval is used as a clean reference to calculate some quality metrics.
from datasets import load_dataset
dataset = load_dataset("yixin/metacloak_celeba_vggface2")
Contact
Contact Us: yixinliucs@gmail.com
danbooru2023-metadata-database
Metadata Database for Danbooru2023
Danbooru 2023 datasets: https://huggingface.co/datasets/nyanko7/danbooru2023
The latest entry of this database is id 7,866,491. Which is newer than nyanko7's dataset.
This dataset contains a sqlite db file which have all the tags and posts metadata in it.
The Peewee ORM config file is provided too, plz check it for more information. (Especially on how I link posts and tags together)
The original data is from the official dump of the posts info.… See the full description on the dataset page: https://huggingface.co/datasets/KBlueLeaf/danbooru2023-metadata-database.
