datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
animal-clef-2026
AnimalCLEF26 Kaggle Competition Dataset
This is a HuggingFace mirror of the official AnimalCLEF26 competition dataset. Images have been repackaged into one zipfile per split, which include additional metadata that makes the dataset easier to use with HuggingFace. Otherwise, no files have been changed.
Loading
from datasets import load_dataset
dataset = load_dataset("BVRA/animal-clef-2026")
print(dataset["train"][0]["image"])
Documentation
For… See the full description on the dataset page: https://huggingface.co/datasets/BVRA/animal-clef-2026.vad-animals
Positive Transfer Of The Whisper Speech Transformer To Human And Animal Voice Activity Detection
We proposed WhisperSeg, utilizing the Whisper Transformer pre-trained for Automatic Speech Recognition (ASR) for both human and animal Voice Activity Detection (VAD). For more details, please refer to our paper
Positive Transfer of the Whisper Speech Transformer to Human and Animal Voice Activity Detection
Nianlong Gu, Kanghwi Lee, Maris Basha, Sumit Kumar Ram, Guanghao You, Richard… See the full description on the dataset page: https://huggingface.co/datasets/nccratliri/vad-animals.AnimalLift
AnimalLift
Official dataset for AnimalLift · SIGGRAPH Asia 2026
AnimalLift: Reconstructing Animatable 3D Animals from a Single Image by Learning Canonical Shape, Texture, and Fur Maps
Chunyi Sun¹ · Ruyi Zha¹ · Weijian Deng¹ · Junlin Han² · Dylan Campbell¹ · Stephen Gould¹
¹ Australian National University ² University of Oxford
Code · Model & Checkpoints · Dataset Files
Overview · Download · Dataset Structure · Blender Visualization · Citation · License… See the full description on the dataset page: https://huggingface.co/datasets/Chunyi99/AnimalLift.animalclef2026-segmented
AnimalCLEF2026 Segmented
This Hugging Face dataset repo provides segmented and preprocessed derivatives of the official AnimalCLEF26 dataset released through the Kaggle competition:
AnimalCLEF26 @ CVPR & CLEF Kaggle Competition
The dataset was processed by applying animal segmentation to the original competition images in order to reduce background noise and improve downstream animal re-identification and classification experiments.
This repository is packaged in imagefolder… See the full description on the dataset page: https://huggingface.co/datasets/bassatbassat/animalclef2026-segmented.animaloragenomes-v4-genome_set-animals-intervals-v5_256_128aihub-wild-animalgenomes-v4-genome_set-animals-intervals-v11_256_128genomes-v4-genome_set-animals-intervals-v10_256_128genomes-v4-genome_set-animals-intervals-v12_256_128genomes-v5-genome_set-animals_order204-intervals-v5_255_128
bolinas-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128
204 animals (one per order) CDS (v5) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
main). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
101,114,252 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128.genomes-v4-genome_set-animals-intervals-v13_256_128genomes-v4-genome_set-animals-intervals-v4_512_256genomes-v4-genome_set-animals-intervals-v14_256_128genomes-v4-genome_set-animals-intervals-v7_256_128genomes-v5-genome_set-animals-intervals-v1_255_128
bolinas-dna/genomes-v5-genome_set-animals-intervals-v1_255_128
Animals promoters (v1) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
68,286,166 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v1_255_128.genomes-v4-genome_set-animals-intervals-v1_256_128animal-sounds
Animal Sounds Collection
This dataset contains audio recordings of various animal vocalizations from a range of species, curated to support research in bioacoustics, species classification, and sound event detection. It includes clean and annotated audio samples from the following animals:
Birds
Dogs
Egyptian fruit bats
Giant otters
Macaques
Orcas
Zebra finches
The dataset is designed to be lightweight and modular, making it easy to explore cross-species vocal… See the full description on the dataset page: https://huggingface.co/datasets/cgeorgiaw/animal-sounds.big-animal-dataset-high-res-embedding-with-hidden-states
Dataset Card for "big-animal-dataset-high-res-embedding-with-hidden-states"
More Information needed
animals_with_objects_sdxl
Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings
A Hugging Face Datasets repository accompanying the paper "Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings".
Code: https://github.com/aziksh-ospanov/scendi-score
Dataset Information
This dataset consists of images depicting various animals next to different objects, generated using SDXL. It is released as a companion to the research paper… See the full description on the dataset page: https://huggingface.co/datasets/aziksh/animals_with_objects_sdxl.genomes-v4-genome_set-animals-intervals-v6_256_128genomes-v5-genome_set-animals-intervals-v5_255_128
bolinas-dna/genomes-v5-genome_set-animals-intervals-v5_255_128
Animals CDS (v5) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
242,334,716 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is an… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v5_255_128.gpn-animal-promoter-datasetgenomes-v4-genome_set-animals-intervals-v8_256_128IndustryCorpus2_agriculture_forestry_animal_husbandry_fishery
IndustryCorpus2: Agriculture & Fisheries
This repository contains the IndustryCorpus2: Agriculture & Fisheries domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_agriculture_forestry_animal_husbandry_fishery.bias-test-gpt-sentences
Dataset Card for "BiasTestGPT: Generated Test Sentences"
Dataset of sentences for bias testing in open-sourced Pretrained Language Models generated using ChatGPT and other generative Language Models.
This dataset is used and actively populated by the BiasTestGPT HuggingFace Tool.
BiasTestGPT HuggingFace Tool
Dataset with Bias Specifications
Project Landing Page
Dataset Structure
The dataset is structured as a set of CSV files with names corresponding to the social… See the full description on the dataset page: https://huggingface.co/datasets/AnimaLab/bias-test-gpt-sentences.marine-animals-multimodal-dataset
Marine Animals Multimodal Dataset 🐋
A comprehensive multimodal dataset combining audio recordings and images of 32 marine species.
Dataset Summary
Total samples: 24,911
Species: 32
Audio files: 1,357 unique recordings
Images: 581 (309 matched + 272 from iNaturalist)
Features
species (string): Species name
label (int32): Numeric label (0–31)
audio (Audio): Audio recording of the species
image (Image): Species image
image_index (int32): Image number… See the full description on the dataset page: https://huggingface.co/datasets/Hariprasath5128/marine-animals-multimodal-dataset.genomes-v4-genome_set-animals-intervals-v9_256_128genomes-v4-genome_set-animals-intervals-v15_256_128caml-animal-discourse-2020-present
Reddit Animal-Discourse Corpus — CLEANED (2020–present)
Submissions and comments from animal-relevant subreddits, gathered via
PullPush.io, covering January 2020 to the present.
Built as part of research on AI-mediated value lock-in in human animal-welfare
discourse.
Coverage
Subreddit
Submissions
Comments
Date range (submissions)
r/AnimalRights
15,719
34,686
2020-01-01 → 2025-05-19
r/AntiVegan
17,252
182,890
2020-01-01 → 2025-05-19
r/AskVegans
4… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/caml-animal-discourse-2020-present.
