datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mixamo-Animations-Characters
Mixamo Animations and Characters
A complete snapshot of the Mixamo library: 2,317 motion clips and
114 rigged characters, exported as binary FBX (FBX 7.7 / fbx7_2019) with per-file metadata.
All animations share one uniform 65-joint mixamorig skeleton, so any clip can drive any
compatible character without remapping.
Use animation_motion/ and character_refined/. The full export contains 2,446 animation
files, but 129 are single-pose assets that carry no motion (Mixamo's *_Pose*… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Mixamo-Animations-Characters.Mixamo-Animations-Characters
Mixamo Animations and Characters
A complete snapshot of the Mixamo library: 2,317 motion clips and
114 rigged characters, exported as binary FBX (FBX 7.7 / fbx7_2019) with per-file metadata.
All animations share one uniform 65-joint mixamorig skeleton, so any clip can drive any
compatible character without remapping.
Use animation_motion/ and character_refined/. The full export contains 2,446 animation
files, but 129 are single-pose assets that carry no motion (Mixamo's *_Pose*… See the full description on the dataset page: https://huggingface.co/datasets/tanish434/Mixamo-Animations-Characters.anime-characters-datasetallaimovies-ai-characters
allaimovies AI characters
3,263 artificial-intelligence characters from 1,884 science-fiction films (1911-2026):
robots, androids, cyborg intelligences, sentient computers, virtual humans and uploaded minds, each
with its kind, how the film presents its gender, whether it helps or opposes the humans, and how
prominent it is. Companion to the allaimovies film dataset from
https://github.com/prateek-0-gupta/allaimovies.
How it was made
For each film in the analysis… See the full description on the dataset page: https://huggingface.co/datasets/prateek-0-gupta/allaimovies-ai-characters.CivitAI-As-CharactersDeduplicated set of CivitAI images as searched by SD XL-derived models that have been described by Llava1.6-34b as Characters.
Each image is a portrait, meaning it's taller than it's wider, and has exactly one face in it. Face bounding boxes are provided.
Character-like description for each image is given by a Llava1.6-34b. Here is an example:
{
"age": "22",
"eyes": "Bright blue, striking",
"face": "Smooth, elegant, with a gentle expression",
"hair": "Long, straight, brown"… See the full description on the dataset page: https://huggingface.co/datasets/kubernetes-bad/CivitAI-As-Characters.gelbooru-characters-enriched
Gelbooru Characters Enriched
This dataset is an enriched, fully-mapped version of Gelbooru character tags. It contains resolved franchise (copyright) associations and core appearance features (core tags) for 263,441 unique characters.
Dataset Details
The dataset maps the original character list to their corresponding copyrights (franchises) and general core attributes. It was constructed using a multi-stage hybrid extraction pipeline:
Regex Extraction: Extracting… See the full description on the dataset page: https://huggingface.co/datasets/cloud19/gelbooru-characters-enriched.big_patent_100k_characters
Sampled Big Patent Dataset
This is a sampled Trelis/big_patent_sample dataset containing rows of data with descriptions shorter than or equal to 100,000 characters in length.
--- Sampled from Trelis/big_patent_sampled ---
Sampled big_patent Dataset
This is a sampled big_patent dataset - sampled down for shorter fine-tunings.
The data is sampled with the aim of providing an even distribution across data lengths. The distribution is quite flat up until 1 million characters… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/big_patent_100k_characters.big_patent_60k_characters
Sampled Trelis/big_patent_sample Dataset
This is a sampled Trelis/big_patent_sample dataset containing rows of data with descriptions shorter than or equal to 60,000 characters in length.
big_patent_60k_to_250k_characters
Sampled Trelis/big_patent_sample Dataset
This is a sampled Trelis/big_patent_sample dataset containing rows of data with descriptions between 60,000 to 250,000 characters in length.
pokemon_omega_ruby_charactersstarwars_charactersgelbooru-characters-onlymerged_characters_tinyllamaPet-Characters
