datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ponys-multilingual-ai-character-consistency-benchmark
Ponys Multilingual AI Character Consistency Benchmark
This repository contains a preregistered test instrument, not collected product
results and not an independent product ranking.
140 fixed test cases across seven locales
four dimensions: persona, register, relationship state, and visual identity
three planned clean-session runs per case
result state: not_collected
publisher: Ponys.ai Research (official first-party research)
official source: https://ponys.ai/
research feeds:… See the full description on the dataset page: https://huggingface.co/datasets/wujoe132/ponys-multilingual-ai-character-consistency-benchmark.allaimovies-ai-characters
allaimovies AI characters
3,263 artificial-intelligence characters from 1,884 science-fiction films (1911-2026):
robots, androids, cyborg intelligences, sentient computers, virtual humans and uploaded minds, each
with its kind, how the film presents its gender, whether it helps or opposes the humans, and how
prominent it is. Companion to the allaimovies film dataset from
https://github.com/prateek-0-gupta/allaimovies.
How it was made
For each film in the analysis… See the full description on the dataset page: https://huggingface.co/datasets/prateek-0-gupta/allaimovies-ai-characters.CivitAI-As-CharactersDeduplicated set of CivitAI images as searched by SD XL-derived models that have been described by Llava1.6-34b as Characters.
Each image is a portrait, meaning it's taller than it's wider, and has exactly one face in it. Face bounding boxes are provided.
Character-like description for each image is given by a Llava1.6-34b. Here is an example:
{
"age": "22",
"eyes": "Bright blue, striking",
"face": "Smooth, elegant, with a gentle expression",
"hair": "Long, straight, brown"… See the full description on the dataset page: https://huggingface.co/datasets/kubernetes-bad/CivitAI-As-Characters.khmer-character-confusions
Khmer Character Confusions
Which Khmer characters people substitute for which, measured from live typing.
When a user is unsure of a spelling they swap a single character and search
again. Each such swap is one row here: the character replaced, the character
tried instead, and how often. No query text appears in this dataset at all —
only character pairs and counts.
The result is an empirical confusability matrix for the Khmer script. It
recovers the expected structure without… See the full description on the dataset page: https://huggingface.co/datasets/seanghay/khmer-character-confusions.character-captions-opusDeduplicated set of character portraits that have been described by Anthropic Claude Opus as characters with stories and visual attributes.
Images obtained from CivitAI by filtering for SD XL-derived models only. Original Stable Diffusion prompt and metadata is also included.
Each image is a portrait, meaning it's taller than it's wider, and has exactly one face in it. Face bounding boxes are provided.
Character-like description for each image is given by Claude Opus. Here is an example:
{… See the full description on the dataset page: https://huggingface.co/datasets/kubernetes-bad/character-captions-opus.africa-synth-aggregates-characterization-nigeria-nigeria
Africa Synth Aggregates Characterization Nigeria Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: infrastructure_transport - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aggregates-characterization-nigeria-nigeria.gelbooru-characters-enriched
Gelbooru Characters Enriched
This dataset is an enriched, fully-mapped version of Gelbooru character tags. It contains resolved franchise (copyright) associations and core appearance features (core tags) for 263,441 unique characters.
Dataset Details
The dataset maps the original character list to their corresponding copyrights (franchises) and general core attributes. It was constructed using a multi-stage hybrid extraction pipeline:
Regex Extraction: Extracting… See the full description on the dataset page: https://huggingface.co/datasets/cloud19/gelbooru-characters-enriched.anomaly-characteristic-layer
The Anomaly Network characteristic layer
A derived dataset over 43,684 first-hand accounts of experiences people
could not explain, drawn from two public archives (NUFORC, 38,663 accounts;
BFRO, 5,021).
Live record: theanomalynetwork.com ·
Dataset page: /data ·
GitHub ·
Zenodo ·
Kaggle
What makes it useful
Every account is coded for which of 63 recurring characteristics it
contains, and every characteristic carries an inverse document frequency.
That IDF column… See the full description on the dataset page: https://huggingface.co/datasets/Rapscallion123/anomaly-characteristic-layer.gelbooru-characters-onlykuzushiji-character-dataset-ogihan-v1
Kuzushiji Character Dataset (Ogihan / Ogi Domain)
This dataset contains single-character Kuzushiji image crops derived from
the Ogihan (小城藩) historical materials, published in a format compatible
with datasets released by CODH (Center for Open Data in the Humanities).
The dataset is designed for:
Kuzushiji OCR
Character-level recognition
Multimodal and vision–language model training
Comparative research with existing CODH datasets
Dataset Structure
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/DimV-Ai/kuzushiji-character-dataset-ogihan-v1.
