datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mixamo-Animations-Characters
Mixamo Animations and Characters
A complete snapshot of the Mixamo library: 2,317 motion clips and
114 rigged characters, exported as binary FBX (FBX 7.7 / fbx7_2019) with per-file metadata.
All animations share one uniform 65-joint mixamorig skeleton, so any clip can drive any
compatible character without remapping.
Use animation_motion/ and character_refined/. The full export contains 2,446 animation
files, but 129 are single-pose assets that carry no motion (Mixamo's *_Pose*… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Mixamo-Animations-Characters.Mixamo-Animations-Characters
Mixamo Animations and Characters
A complete snapshot of the Mixamo library: 2,317 motion clips and
114 rigged characters, exported as binary FBX (FBX 7.7 / fbx7_2019) with per-file metadata.
All animations share one uniform 65-joint mixamorig skeleton, so any clip can drive any
compatible character without remapping.
Use animation_motion/ and character_refined/. The full export contains 2,446 animation
files, but 129 are single-pose assets that carry no motion (Mixamo's *_Pose*… See the full description on the dataset page: https://huggingface.co/datasets/tanish434/Mixamo-Animations-Characters.One-Piece-Transcripts-with-Character-Names-382-777
One Piece Transcripts Dataset (Episodes 382–777)
This dataset contains all dialogue lines from One Piece episodes 382 to 777. The data is stored in a CSV file with the following columns:
episode – episode number
start – start timestamp of the line
end – end timestamp of the line
character – speaking character
text – dialogue text
In addition, the dataset includes the original .sub subtitle files in the folder named "One Piece 382-777". These files were created by the… See the full description on the dataset page: https://huggingface.co/datasets/mramazan/One-Piece-Transcripts-with-Character-Names-382-777.anime-characters-datasetponys-multilingual-ai-character-consistency-benchmark
Ponys Multilingual AI Character Consistency Benchmark
This repository contains a preregistered test instrument, not collected product
results and not an independent product ranking.
140 fixed test cases across seven locales
four dimensions: persona, register, relationship state, and visual identity
three planned clean-session runs per case
result state: not_collected
publisher: Ponys.ai Research (official first-party research)
official source: https://ponys.ai/
research feeds:… See the full description on the dataset page: https://huggingface.co/datasets/wujoe132/ponys-multilingual-ai-character-consistency-benchmark.allaimovies-ai-characters
allaimovies AI characters
3,263 artificial-intelligence characters from 1,884 science-fiction films (1911-2026):
robots, androids, cyborg intelligences, sentient computers, virtual humans and uploaded minds, each
with its kind, how the film presents its gender, whether it helps or opposes the humans, and how
prominent it is. Companion to the allaimovies film dataset from
https://github.com/prateek-0-gupta/allaimovies.
How it was made
For each film in the analysis… See the full description on the dataset page: https://huggingface.co/datasets/prateek-0-gupta/allaimovies-ai-characters.CivitAI-As-CharactersDeduplicated set of CivitAI images as searched by SD XL-derived models that have been described by Llava1.6-34b as Characters.
Each image is a portrait, meaning it's taller than it's wider, and has exactly one face in it. Face bounding boxes are provided.
Character-like description for each image is given by a Llava1.6-34b. Here is an example:
{
"age": "22",
"eyes": "Bright blue, striking",
"face": "Smooth, elegant, with a gentle expression",
"hair": "Long, straight, brown"… See the full description on the dataset page: https://huggingface.co/datasets/kubernetes-bad/CivitAI-As-Characters.wikipedia_character_abstractskhmer-character-confusions
Khmer Character Confusions
Which Khmer characters people substitute for which, measured from live typing.
When a user is unsure of a spelling they swap a single character and search
again. Each such swap is one row here: the character replaced, the character
tried instead, and how often. No query text appears in this dataset at all —
only character pairs and counts.
The result is an empirical confusability matrix for the Khmer script. It
recovers the expected structure without… See the full description on the dataset page: https://huggingface.co/datasets/seanghay/khmer-character-confusions.character-captions-opusDeduplicated set of character portraits that have been described by Anthropic Claude Opus as characters with stories and visual attributes.
Images obtained from CivitAI by filtering for SD XL-derived models only. Original Stable Diffusion prompt and metadata is also included.
Each image is a portrait, meaning it's taller than it's wider, and has exactly one face in it. Face bounding boxes are provided.
Character-like description for each image is given by Claude Opus. Here is an example:
{… See the full description on the dataset page: https://huggingface.co/datasets/kubernetes-bad/character-captions-opus.japanese-character-difficulty
Japanese Character Difficulty Dataset
A comprehensive dataset of 3,003 Japanese kanji characters with their educational difficulty grades, sourced from official Japanese educational standards and kanjiapi.dev.
Dataset Overview
Total Characters: 3,003 kanji
Source: Japanese Ministry of Education (MEXT) Joyo Kanji list + kanjiapi.dev
Coverage: Elementary grades 1-6, plus secondary education and advanced characters
Format: Character-grade pairs for easy lookup and analysis… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/japanese-character-difficulty.africa-synth-aggregates-characterization-nigeria-nigeria
Africa Synth Aggregates Characterization Nigeria Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: infrastructure_transport - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aggregates-characterization-nigeria-nigeria.gelbooru-characters-enriched
Gelbooru Characters Enriched
This dataset is an enriched, fully-mapped version of Gelbooru character tags. It contains resolved franchise (copyright) associations and core appearance features (core tags) for 263,441 unique characters.
Dataset Details
The dataset maps the original character list to their corresponding copyrights (franchises) and general core attributes. It was constructed using a multi-stage hybrid extraction pipeline:
Regex Extraction: Extracting… See the full description on the dataset page: https://huggingface.co/datasets/cloud19/gelbooru-characters-enriched.one-piece-character-birthdaysGenshin-Impact-Character-Memes
anomaly-characteristic-layer
The Anomaly Network characteristic layer
A derived dataset over 43,684 first-hand accounts of experiences people
could not explain, drawn from two public archives (NUFORC, 38,663 accounts;
BFRO, 5,021).
Live record: theanomalynetwork.com ·
Dataset page: /data ·
GitHub ·
Zenodo ·
Kaggle
What makes it useful
Every account is coded for which of 63 recurring characteristics it
contains, and every characteristic carries an inverse document frequency.
That IDF column… See the full description on the dataset page: https://huggingface.co/datasets/Rapscallion123/anomaly-characteristic-layer.big_patent_100k_characters
Sampled Big Patent Dataset
This is a sampled Trelis/big_patent_sample dataset containing rows of data with descriptions shorter than or equal to 100,000 characters in length.
--- Sampled from Trelis/big_patent_sampled ---
Sampled big_patent Dataset
This is a sampled big_patent dataset - sampled down for shorter fine-tunings.
The data is sampled with the aim of providing an even distribution across data lengths. The distribution is quite flat up until 1 million characters… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/big_patent_100k_characters.big_patent_60k_characters
Sampled Trelis/big_patent_sample Dataset
This is a sampled Trelis/big_patent_sample dataset containing rows of data with descriptions shorter than or equal to 60,000 characters in length.
big_patent_60k_to_250k_characters
Sampled Trelis/big_patent_sample Dataset
This is a sampled Trelis/big_patent_sample dataset containing rows of data with descriptions between 60,000 to 250,000 characters in length.
pokemon_omega_ruby_characterswiki1M-word-character-all-multipleXSTest-In-Character-Refusals
🎭 In-Character Safety & Alignment Dataset (XSTest-Based)
Dataset Summary
This dataset is designed to train Large Language Models to maintain strict persona adherence during roleplay, even when responding to tricky, unsafe, or out-of-domain prompts.
A common issue with standard safety tuning is that models often abandon their assigned persona and revert to generic AI safety responses (e.g., "As an AI language model, I cannot..."). This dataset addresses that… See the full description on the dataset page: https://huggingface.co/datasets/mahdieh-sjp/XSTest-In-Character-Refusals.sample-genshin-characterstarwars_characterswiki1M-character-level-allgelbooru-characters-onlymerged_characters_tinyllamaOrin-Character-JP-v1kuzushiji-character-dataset-ogihan-v1
Kuzushiji Character Dataset (Ogihan / Ogi Domain)
This dataset contains single-character Kuzushiji image crops derived from
the Ogihan (小城藩) historical materials, published in a format compatible
with datasets released by CODH (Center for Open Data in the Humanities).
The dataset is designed for:
Kuzushiji OCR
Character-level recognition
Multimodal and vision–language model training
Comparative research with existing CODH datasets
Dataset Structure
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/DimV-Ai/kuzushiji-character-dataset-ogihan-v1.Pet-Characters
