datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vad-animals
Positive Transfer Of The Whisper Speech Transformer To Human And Animal Voice Activity Detection
We proposed WhisperSeg, utilizing the Whisper Transformer pre-trained for Automatic Speech Recognition (ASR) for both human and animal Voice Activity Detection (VAD). For more details, please refer to our paper
Positive Transfer of the Whisper Speech Transformer to Human and Animal Voice Activity Detection
Nianlong Gu, Kanghwi Lee, Maris Basha, Sumit Kumar Ram, Guanghao You, Richard… See the full description on the dataset page: https://huggingface.co/datasets/nccratliri/vad-animals.animal-sounds
Animal Sounds Collection
This dataset contains audio recordings of various animal vocalizations from a range of species, curated to support research in bioacoustics, species classification, and sound event detection. It includes clean and annotated audio samples from the following animals:
Birds
Dogs
Egyptian fruit bats
Giant otters
Macaques
Orcas
Zebra finches
The dataset is designed to be lightweight and modular, making it easy to explore cross-species vocal… See the full description on the dataset page: https://huggingface.co/datasets/cgeorgiaw/animal-sounds.marine-animals-multimodal-dataset
Marine Animals Multimodal Dataset 🐋
A comprehensive multimodal dataset combining audio recordings and images of 32 marine species.
Dataset Summary
Total samples: 24,911
Species: 32
Audio files: 1,357 unique recordings
Images: 581 (309 matched + 272 from iNaturalist)
Features
species (string): Species name
label (int32): Numeric label (0–31)
audio (Audio): Audio recording of the species
image (Image): Species image
image_index (int32): Image number… See the full description on the dataset page: https://huggingface.co/datasets/Hariprasath5128/marine-animals-multimodal-dataset.Human-Animal-Cartoon
Human-Animal-Cartoon dataset
Our Human-Animal-Cartoon (HAC) dataset consists of seven actions (‘sleeping’, ‘watching tv’, ‘eating’, ‘drinking’, ‘swimming’, ‘running’, and ‘opening door’) performed by humans, animals, and cartoon figures, forming three different domains. We collect 3381 video clips from the internet with around 1000 for each domain and provide three modalities in our dataset: video, audio, and pre-computed optical flow.
The dataset can be used for Multi-modal Domain… See the full description on the dataset page: https://huggingface.co/datasets/hdong51/Human-Animal-Cartoon.Animal-Sound-Instructions
Animal Sound Instructions
We gathered from,
Birds, birdclef-2021
Insecta, christopher/birdclef-2025
Amphibia, christopher/birdclef-2025
Mammalia, christopher/birdclef-2025
We use Qwen/Qwen2.5-72B-Instruct to generate the answers based on the metadata.
how to prepare the dataset
huggingface-cli download \
mesolitica/Animal-Sound-Instructions \
--include "*.zip" \
--repo-type "dataset" \
--local-dir './'
wget… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Animal-Sound-Instructions.marine-animals-multimodalEnvironmentalSoundClassification_ESC50-Animals
Dataset Card for "environmental_sound_classification_animals_ESC50"
More Information needed
animalAudioSet-Strong-AnimalsHuman-Animal-Cartoon-PC-VAanimalclap-dataset
AnimalCLAP
AnimalCLAP: Taxonomy-Aware Language-Audio Pretraining for Species Recognition and Trait InferenceICASSP 2026
AuthorsRisa Shinoda, Kaede Shiohara, Nakamasa Inoue, Hiroaki Santo, Fumio Okura
Overview
This dataset contains 701,020 animal sound recordings collected from:
iNaturalist
Xeno-Canto
Splits
HF Split
Original Split
Description
train
train
Training data (URL only)
validation
test
Validation data (URL only)
test
zero_shot… See the full description on the dataset page: https://huggingface.co/datasets/risashinoda/animalclap-dataset.animalspeak-pseudovox
AnimalSpeak Pseudovox Train-Unseen
This dataset contains the train-unseen split of AnimalSpeak Pseudovox. Each
example is a short, silence-trimmed, single-vocalization WAV clip plus compact
per-clip metadata. It does not include generated conversations, captions, QA
pairs, or MCQ answers.
Rows: 346,907
Shards: 18
Maximum rows per shard: 20,000
Files
data-20k/train-*.tar: WebDataset-style shards containing
audio/<audio_name> WAV entries.
metadata.parquet: one row per… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/animalspeak-pseudovox.AnimalQA
Dataset Card for SAKURA-AnimalQA
This dataset contains the audio and the single/multi-hop questions/answers of the animal track of the SAKURA benchmark from Interspeech 2025 paper, "SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information".
The fields of the dataset are:
file: The filename of the audio files.
audio: The audio recordings.
attribute_label: The attribute labels (i.e., the kinds of animal making the sounds) of the audio… See the full description on the dataset page: https://huggingface.co/datasets/SLLM-multi-hop/AnimalQA.EnvironmentalSoundClassification_ESC50-Animals_TTSanimal_classificationHuman-Animal-CartoonAnimalAudioA dataset to fine-tune the AudioLDM Audio Generation Model
AnimalQA-gemini-1.5-pro-caption
Dataset Card for "AnimalQA-gemini-1.5-pro-caption"
More Information needed
AnimalClassification_WaveSource-TestAnimalQA-gemini-1.5-flash-fix
Dataset Card for "AnimalQA-gemini-1.5-flash-fix"
More Information needed
AnimalQA-gemini-1.5-pro
Dataset Card for "AnimalQA-gemini-1.5-pro"
More Information needed
AnimalQA-gemini-1.5-flash-caption
Dataset Card for "AnimalQA-gemini-1.5-flash-caption"
More Information needed
AnimalQA-LTUAS-caption
Dataset Card for "AnimalQA-LTUAS-caption"
More Information needed
AnimalQA-LTUAS
Dataset Card for "AnimalQA-LTUAS"
More Information needed
audio-alpaca-animalsAnimalQA-gemini-1.5-flash
Dataset Card for "AnimalQA-gemini-1.5-flash"
More Information needed
AnimalQA-GPT4o
Dataset Card for "AnimalQA-GPT4o"
More Information needed
AnimalQA-GPT4o-caption
Dataset Card for "AnimalQA-GPT4o-caption"
More Information needed
AnimalQA-Best-Balanced-DistractorsAnimaux_RAW_AUDIO
