datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vad-animals
Positive Transfer Of The Whisper Speech Transformer To Human And Animal Voice Activity Detection
We proposed WhisperSeg, utilizing the Whisper Transformer pre-trained for Automatic Speech Recognition (ASR) for both human and animal Voice Activity Detection (VAD). For more details, please refer to our paper
Positive Transfer of the Whisper Speech Transformer to Human and Animal Voice Activity Detection
Nianlong Gu, Kanghwi Lee, Maris Basha, Sumit Kumar Ram, Guanghao You, Richard… See the full description on the dataset page: https://huggingface.co/datasets/nccratliri/vad-animals.animal-sounds
Animal Sounds Collection
This dataset contains audio recordings of various animal vocalizations from a range of species, curated to support research in bioacoustics, species classification, and sound event detection. It includes clean and annotated audio samples from the following animals:
Birds
Dogs
Egyptian fruit bats
Giant otters
Macaques
Orcas
Zebra finches
The dataset is designed to be lightweight and modular, making it easy to explore cross-species vocal… See the full description on the dataset page: https://huggingface.co/datasets/cgeorgiaw/animal-sounds.animals_with_objects_sdxl
Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings
A Hugging Face Datasets repository accompanying the paper "Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings".
Code: https://github.com/aziksh-ospanov/scendi-score
Dataset Information
This dataset consists of images depicting various animals next to different objects, generated using SDXL. It is released as a companion to the research paper… See the full description on the dataset page: https://huggingface.co/datasets/aziksh/animals_with_objects_sdxl.genomes-v4-genome_set-animals-intervals-v5_256_128genomes-v4-genome_set-animals-intervals-v4_512_256genomes-v4-genome_set-animals-intervals-v11_256_128genomes-v5-genome_set-animals-intervals-v1_255_128
bolinas-dna/genomes-v5-genome_set-animals-intervals-v1_255_128
Animals promoters (v1) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
68,286,166 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v1_255_128.genomes-v4-genome_set-animals-intervals-v10_256_128genomes-v4-genome_set-animals-intervals-v12_256_128genomes-v4-genome_set-animals-intervals-v13_256_128genomes-v4-genome_set-animals-intervals-v14_256_128genomes-v4-genome_set-animals-intervals-v7_256_128genomes-v5-genome_set-animals_order204-intervals-v5_255_128
bolinas-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128
204 animals (one per order) CDS (v5) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
main). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
101,114,252 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128.genomes-v4-genome_set-animals-intervals-v1_256_128genomes-v4-genome_set-animals-intervals-v6_256_128genomes-v4-genome_set-animals-intervals-v8_256_128genomes-v5-genome_set-animals-intervals-v5_255_128
bolinas-dna/genomes-v5-genome_set-animals-intervals-v5_255_128
Animals CDS (v5) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
242,334,716 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is an… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v5_255_128.genomes-v4-genome_set-animals-intervals-v9_256_128genomes-v4-genome_set-animals-intervals-v15_256_128genomes-v5-genome_set-animals-intervals-v15_255_128
bolinas-dna/genomes-v5-genome_set-animals-intervals-v15_255_128
Animals downstream-of-CDS (v15) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
20,501,856 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v15_255_128.PUG_Animals
PUG Animals
The PUG: Animals dataset contains 215,040 pre-rendered images based on Unreal-Engine using 70 animal assets, 64 environments, 3 sizes, 4 textures, under 4 camera orientations.
It was designed with the intent to create a dataset with variation factors available. Inspired by research on out-of-distribution generalization, PUG: Animals allows one to precisely control distribution shifts between training and testing which can provide better insight on how a deep neural… See the full description on the dataset page: https://huggingface.co/datasets/facebook/PUG_Animals.marine-animals-multimodal-dataset
Marine Animals Multimodal Dataset 🐋
A comprehensive multimodal dataset combining audio recordings and images of 32 marine species.
Dataset Summary
Total samples: 24,911
Species: 32
Audio files: 1,357 unique recordings
Images: 581 (309 matched + 272 from iNaturalist)
Features
species (string): Species name
label (int32): Numeric label (0–31)
audio (Audio): Audio recording of the species
image (Image): Species image
image_index (int32): Image number… See the full description on the dataset page: https://huggingface.co/datasets/Hariprasath5128/marine-animals-multimodal-dataset.genomes-v3-genome_set-animals-intervals-v3_512_256genomes-v3-genome_set-animals-intervals-v2_512_256genomes-v3-genome_set-animals-intervals-v1_512_256
Animal promoters
genomes-v2-genome_set-animals-intervals-v1_512_256animals-cds-proj-v1-vert125
Animal CDS via human→animal projection — vertebrates (Chordata)
Human protein-coding (CDS) 255 bp windows projected onto 125 vertebrate species (Chordata, one per order) by mmseqs2 nucleotide local alignment — the projection ("Arm B") side of the issue #353 CDS projection-vs-annotation experiment.
How it was built
The human v5 CDS intervals (Ensembl coding exons, filtered 20–10,000 bp, +20 bp splice flank, expanded ≥256 bp) are tiled into 255 bp windows (269,866… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/animals-cds-proj-v1-vert125.animals-cds-proj-v1-all204
Animal CDS via human→animal projection — all animal orders
Human protein-coding (CDS) 255 bp windows projected onto 204 animal species (one annotated genome per Metazoan order) by mmseqs2 nucleotide local alignment — the projection ("Arm B") side of the issue #353 CDS projection-vs-annotation experiment.
How it was built
The human v5 CDS intervals (Ensembl coding exons, filtered 20–10,000 bp, +20 bp splice flank, expanded ≥256 bp) are tiled into 255 bp windows… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/animals-cds-proj-v1-all204.golden-data-animals
Golden Data: Washin Animal Village
1554 high-quality animal images.
animals-10
