datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mixamo-Animations-Characters
Mixamo Animations and Characters
A complete snapshot of the Mixamo library: 2,317 motion clips and
114 rigged characters, exported as binary FBX (FBX 7.7 / fbx7_2019) with per-file metadata.
All animations share one uniform 65-joint mixamorig skeleton, so any clip can drive any
compatible character without remapping.
Use animation_motion/ and character_refined/. The full export contains 2,446 animation
files, but 129 are single-pose assets that carry no motion (Mixamo's *_Pose*… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Mixamo-Animations-Characters.Mixamo-Animations-Characters
Mixamo Animations and Characters
A complete snapshot of the Mixamo library: 2,317 motion clips and
114 rigged characters, exported as binary FBX (FBX 7.7 / fbx7_2019) with per-file metadata.
All animations share one uniform 65-joint mixamorig skeleton, so any clip can drive any
compatible character without remapping.
Use animation_motion/ and character_refined/. The full export contains 2,446 animation
files, but 129 are single-pose assets that carry no motion (Mixamo's *_Pose*… See the full description on the dataset page: https://huggingface.co/datasets/tanish434/Mixamo-Animations-Characters.anima
ANIMA: Animal Norms In Moral Assessment
ANIMA is a dataset of 26 prompts and 13 ethical dimensions used to evaluate the quality of a model's moral reasoning about animal welfare. It is the dataset behind the inspect_evals/anima eval and is introduced in Alignment Midtraining for Animals (Brazilek & Tidmarsh, 2026).
Renamed from AHB. This dataset was previously published as the Animal Harm Benchmark (AHB) at sentientfutures/ahb. It was renamed to ANIMA in May 2026 to… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/anima.animal-clef-2026
AnimalCLEF26 Kaggle Competition Dataset
This is a HuggingFace mirror of the official AnimalCLEF26 competition dataset. Images have been repackaged into one zipfile per split, which include additional metadata that makes the dataset easier to use with HuggingFace. Otherwise, no files have been changed.
Loading
from datasets import load_dataset
dataset = load_dataset("BVRA/animal-clef-2026")
print(dataset["train"][0]["image"])
Documentation
For… See the full description on the dataset page: https://huggingface.co/datasets/BVRA/animal-clef-2026.animal-sounds
Animal Sounds Collection
This dataset contains audio recordings of various animal vocalizations from a range of species, curated to support research in bioacoustics, species classification, and sound event detection. It includes clean and annotated audio samples from the following animals:
Birds
Dogs
Egyptian fruit bats
Giant otters
Macaques
Orcas
Zebra finches
The dataset is designed to be lightweight and modular, making it easy to explore cross-species vocal… See the full description on the dataset page: https://huggingface.co/datasets/cgeorgiaw/animal-sounds.TA2.0_animation_dataset
TA2.0 Animation Dataset for Talking Avatars
11,255 AI-generated anime-style talking avatar video clips (≈30.3 hours) with per-clip quality metadata.
Each clip shows a single animated character speaking, rendered at 1280×1280 @ 25 fps (H.264 video + AAC audio, ~10 s per clip). Clips were generated with LongCat-Video-Avatar (Meituan LongCat Team), an audio-driven avatar video generation model, conditioned on reference character images and driven by speech audio in 9 languages —… See the full description on the dataset page: https://huggingface.co/datasets/neosapience/TA2.0_animation_dataset.Objaverse-XL-Rigged-Animated
Objaverse-XL Rigged & Animated Subset
Every asset here carries both a skeleton and at least one animation clip, selected from
Objaverse / Objaverse-XL. Rigs range from 3 to 344 joints and
span characters as well as articulated rigid objects.
Objaverse-XL indexes over 10 million objects, but only a small fraction carry a usable rig and
motion on it. This subset isolates that fraction: every file was checked to contain at least one
skin with joints and at least one animation clip… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Objaverse-XL-Rigged-Animated.Diffusion4D-Animated-Raw
Diffusion4D Animated Assets
This dataset provides animated 3D assets referenced by Diffusion4D and
Objaverse-XL in a directly browsable format. The default split contains 67,988
rows. Each row includes metadata, a preview image, and a short preview video so
that assets can be inspected in the Hugging Face Data Studio without first
downloading the original 3D file.
The repository also mirrors available raw assets and keeps their original
source links and hashes. The current… See the full description on the dataset page: https://huggingface.co/datasets/DenisKochetov/Diffusion4D-Animated-Raw.wan22-animate-3k-opensource-data
Wan2.2 Animate Open Dataset Pack
This dataset repo stores the complete datasets/ directory used for the Wan2.2 TI2V 5B + One-to-All animate experiment.
The original tree contains more than 10,000 files in one directory, which Hugging Face git repositories reject as raw files. Therefore the dataset is stored as split tar shards.
Restore:
cat datasets.tar.part-* | tar -xf -
sha256sum -c SHA256SUMS
After extraction, the restored tree contains:… See the full description on the dataset page: https://huggingface.co/datasets/simbahuang/wan22-animate-3k-opensource-data.genomes-v4-genome_set-animals-intervals-v5_256_128genomes-v4-genome_set-animals-intervals-v4_512_256genomes-v4-genome_set-animals-intervals-v11_256_128genomes-v5-genome_set-animals-intervals-v1_255_128
bolinas-dna/genomes-v5-genome_set-animals-intervals-v1_255_128
Animals promoters (v1) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
68,286,166 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v1_255_128.genomes-v4-genome_set-animals-intervals-v10_256_128genomes-v4-genome_set-animals-intervals-v12_256_128genomes-v4-genome_set-animals-intervals-v13_256_128genomes-v4-genome_set-animals-intervals-v14_256_128gpn-animal-promoter-datasetgenomes-v5-genome_set-animals_order204-intervals-v5_255_128
bolinas-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128
204 animals (one per order) CDS (v5) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
main). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
101,114,252 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128.genomes-v4-genome_set-animals-intervals-v1_256_128bias-test-gpt-sentences
Dataset Card for "BiasTestGPT: Generated Test Sentences"
Dataset of sentences for bias testing in open-sourced Pretrained Language Models generated using ChatGPT and other generative Language Models.
This dataset is used and actively populated by the BiasTestGPT HuggingFace Tool.
BiasTestGPT HuggingFace Tool
Dataset with Bias Specifications
Project Landing Page
Dataset Structure
The dataset is structured as a set of CSV files with names corresponding to the social… See the full description on the dataset page: https://huggingface.co/datasets/AnimaLab/bias-test-gpt-sentences.mini-reachy-animation
Reachy Mini Animation Dataset
Multi-view renders of 85 emotional animations performed by the Reachy Mini
robot, paired with the full robot joint state for every single frame.
Source of the animations. The emotional animations rendered here come from the
official pollen-robotics/reachy-mini-emotions-library
dataset by Pollen Robotics. This dataset re-renders those emotions from 12 camera
angles (with 3 background variants) and pairs every frame with the robot's joint state.… See the full description on the dataset page: https://huggingface.co/datasets/BastienATOS/mini-reachy-animation.genomes-v4-genome_set-animals-intervals-v6_256_128genomes-v4-genome_set-animals-intervals-v8_256_128genomes-v5-genome_set-animals-intervals-v5_255_128
bolinas-dna/genomes-v5-genome_set-animals-intervals-v5_255_128
Animals CDS (v5) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
242,334,716 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is an… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v5_255_128.genomes-v4-genome_set-animals-intervals-v9_256_128Objaverse-XL-Rigged-Animated-Renders
Objaverse-XL Rigged & Animated — Renders
Visual companion to
Linzhan/Objaverse-XL-Rigged-Animated,
which holds the 7,373 rigged-and-animated GLB assets themselves. This repository holds only what
was rendered from them: a four-view video of every animation clip, and a rest-pose grid per asset.
They live apart from the assets because they are bulky and numerous — 10,355 clip folders — while
the asset repo stays a compact 7,373 GLBs plus two tables. Nothing here is needed to use… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Objaverse-XL-Rigged-Animated-Renders.genomes-v4-genome_set-animals-intervals-v15_256_128animal_dataset_activations
Animal Dataset Activations
This dataset contains neural network activations captured from various models processing the animal_dataset.
Dataset Structure
The dataset is organized by model (subset) and layer (split):
Model (subset): Each model has its own directory
Layer (split): Each layer within a model has its own directory with parquet files
Files are organized as: {model_name}/{layer_name}/*.parquet
Loading the Dataset
The dataset uses HuggingFace's… See the full description on the dataset page: https://huggingface.co/datasets/nirmalendu01/animal_dataset_activations.genomes-v5-genome_set-animals-intervals-v15_255_128
bolinas-dna/genomes-v5-genome_set-animals-intervals-v15_255_128
Animals downstream-of-CDS (v15) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
20,501,856 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v15_255_128.
