datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dynamic_earthnetDynamic EarthNet dataset redistributed from https://mediatum.ub.tum.de/1650201 and https://cvg.cit.tum.de/webshare/u/toker/dynnet_training_splits/ under a common tarball for simpler download speeds.
Individual zip files were replaced with tarballs instead.
In the mediatum server version the following directories have the wrong name compared to the given split txt files:
/labels/5111_4560_13_38S
/labels/6204_3495_13_46N
/labels/7026_3201_13_52N
/labels/7367_5050_13_54S
/labels/2459_4406_13_19S… See the full description on the dataset page: https://huggingface.co/datasets/torchgeo/dynamic_earthnet.LaSeRSanimalspeak-pseudovox
AnimalSpeak Pseudovox Train-Unseen
This dataset contains the train-unseen split of AnimalSpeak Pseudovox. Each
example is a short, silence-trimmed, single-vocalization WAV clip plus compact
per-clip metadata. It does not include generated conversations, captions, QA
pairs, or MCQ answers.
Rows: 346,907
Shards: 18
Maximum rows per shard: 20,000
Files
data-20k/train-*.tar: WebDataset-style shards containing
audio/<audio_name> WAV entries.
metadata.parquet: one row per… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/animalspeak-pseudovox.DE-Dataset
