datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imagenet-aThe ImageNet-A dataset contains 7,500 natural adversarial examples.
Source: https://github.com/hendrycks/natural-adv-examples.Also see the ImageNet-C and ImageNet-P datasets at https://github.com/hendrycks/robustness
@article{hendrycks2019nae, title={Natural Adversarial Examples}, author={Dan Hendrycks and Kevin Zhao and Steven Basart and Jacob Steinhardt and Dawn Song}, journal={arXiv preprint arXiv:1907.07174}, year={2019}}
There are 200 classes we consider. The WordNet ID and a… See the full description on the dataset page: https://huggingface.co/datasets/barkermrl/imagenet-a.Barkopedia-Dog-Vocal-Detection
🐾 Dog Vocal Detection
This dataset is curated from internet videos to support research in dog vocalization detection using both weak and strong supervision.
It contains approximately 7,500 seconds of strongly labeled training audio
Over 9,000 seconds of weakly labeled clips sourced from AudioSet are included.
The dataset also provides 24 hours of unlabeled audio clips from our own collection.
To simulate realistic conditions, some clips feature dogs present without barking… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/Barkopedia-Dog-Vocal-Detection.Barkopedia_Dog_Sex_Classification_Dataset
📦 Dataset Description
This dataset is part of the Barkopedia Challenge: https://uta-acl2.github.io/barkopedia.html
Check training data on Hugging Face:
👉 ArlingtonCL2/Barkopedia_Dog_Sex_Classification_Dataset
This challenge provides a dataset of labeled dog bark audio clips:
29,345 total clips of vocalizations from 156 individual dogs across 5 breeds:
Shiba Inu
Husky
Chihuahua
German Shepherd
Pitbull
Training set: 26,895 clips
13,567 female13,328 male
Test set: 2,450… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/Barkopedia_Dog_Sex_Classification_Dataset.Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET
Dataset
Check Training Data here: ArlingtonCL2/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET
Dataset Description
This dataset is for Dog Age Group Classification and contains dog bark audio clips. The data is split into training, public test, and private test sets.
Training set: 17888 audio clips.
Test set: 4920 audio clips, further divided into:
Test Public (~40%): 1966 audio clips for live leaderboard updates.
Test Private (~60%): 2954 audio clips for final evaluation.
You will… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET.BarkVN-50
Dataset Card for BarkVN-50: Tree Species Identification from Bark Texture
This is a FiftyOne dataset with 5578 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/BarkVN-50")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/BarkVN-50.womenBarkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET
Dataset
Check Training Data here: ArlingtonCL2/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET
Dataset Description
This dataset is for Dog Age Group Classification and contains dog bark audio clips. The data is split into training, public test, and private test sets.
Training set: 17888 audio clips.
Test set: 4920 audio clips, further divided into:
Test Public (~40%): 1966 audio clips for live leaderboard updates.
Test Private (~60%): 2954 audio clips for final evaluation.
You… See the full description on the dataset page: https://huggingface.co/datasets/hlx1021/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET.Barkopedia_DOG_BREED_CLASSIFICATION_DATASET
📦 Dataset Description
This dataset is part of the Barkopedia Challenge
🔗 https://uta-acl2.github.io/barkopedia.html
Check Training Data here:👉 ArlingtonCL2/Barkopedia_DOG_BREED_CLASSIFICATION_DATASET
This dataset contains 29,347 audio clips of dog barks labeled by dog breed.
The audio samples come from 156 individual dogs across 5 dog breeds:
shiba inu
husky
chihuahua
german shepherd
pitbull
📊 Per-Breed Summary
Breed
Train
Public TestPrivate Test
Test… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/Barkopedia_DOG_BREED_CLASSIFICATION_DATASET.barkai_compendium
Barkai Compendium
This collects the ChEC-seq data from the following GEO series:
GSE179430
GSE209631
GSE222268
The metadata for each is parsed out from the SraRunTable, or in the case of GSE222268,
the NCBI series matrix file (the genotype isn't in the SraRunTable)
The Barkai lab refers to this set as their
binding compendium.
The genotypes for GSE222268 are not clear enough to me currently to parse well.
Accessing Data
The examples below require the
HuggingFace Hub… See the full description on the dataset page: https://huggingface.co/datasets/BrentLab/barkai_compendium.BarkopediaDogEmotionClassification_Data
EmotionalCanines: A Dataset for Analysis of Arousal and Valence in Dog Vocalization
Paper: https://dl.acm.org/doi/10.1145/3746027.3758286
Full Dataset: https://github.com/tmdang1101/EmotionalCanines
🔖 Labels
Each clip is annotated with one arousal (Low, Medium, High) and one valence (Negative, Neutral, Positive) category.
Note Regarding Data Split
The sets do not contain the same dogs.
Citation
If you use this dataset in your… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/BarkopediaDogEmotionClassification_Data.audio-course-bark-samplesbark-ambrosia-beetle-benchmark
Bark and Ambrosia Beetle Detection Benchmark
Version 2.0.1 · 14,491 images · 175 species · 21 tribes · 70 genera · COCO detection format
A specimen-disjoint, species-level object-detection benchmark for bark and ambrosia
beetles (Coleoptera: Curculionidae: Scolytinae and Platypodinae), derived from the
Bark and Ambrosia Gallery (https://barkandambrosiagallery.org/). Species
determinations are made or reviewed by taxonomists; individual specimens carry
bounding boxes.
This Zenodo… See the full description on the dataset page: https://huggingface.co/datasets/IBBI-bio/bark-ambrosia-beetle-benchmark.Barkopedia_Dog_Act_Env
🐶 Barkopedia Challenge Dataset
🔗 Barkopedia Website
📦 Dataset Description
This challenge provides a labeled dataset of dog bark audio clips for understanding activity and environment from sound.
📁 Current Release
Training Set
Located in the train/ folder
Includes:
split1.zip and split2.zip — each contains a portion of the audio files
train_label.csv — contains labels for all training clips
Total: 12,480 training audio clips
Test Set
To be released in… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/Barkopedia_Dog_Act_Env.details_Slim205__Barka-9b-it_v2
Dataset Card for Evaluation run of Slim205/Barka-9b-it
Dataset automatically created during the evaluation run of model Slim205/Barka-9b-it.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Slim205__Barka-9b-it_v2.Barkopedia_DOG_BREED_CLASSIFICATION_DATASET
📦 Dataset Description
This dataset is part of the Barkopedia Challenge
🔗 https://uta-acl2.github.io/barkopedia.html
Check Training Data here:👉 ArlingtonCL2/Barkopedia_DOG_BREED_CLASSIFICATION_DATASET
This dataset contains 29,347 audio clips of dog barks labeled by dog breed.
The audio samples come from 156 individual dogs across 5 dog breeds:
shiba inu
husky
chihuahua
german shepherd
pitbull
📊 Per-Breed Summary
Breed
Train
Public Test
Private Test… See the full description on the dataset page: https://huggingface.co/datasets/PuneettArora/Barkopedia_DOG_BREED_CLASSIFICATION_DATASET.slurp_synthetic_barkdetails_Slim205__Barka-2b-it_v2
Dataset Card for Evaluation run of Slim205/Barka-2b-it
Dataset automatically created during the evaluation run of model Slim205/Barka-2b-it.
The dataset is composed of 114 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Slim205__Barka-2b-it_v2.mnist-cSource: https://github.com/google-research/mnist-c
MNIST-C
This repository contains the source code used to create the MNIST-C dataset, a
corrupted MNIST benchmark for testing out-of-distribution robustness of computer
vision models.
Please see our full paper https://arxiv.org/abs/1906.02337 for more details.
Dataset
The static dataset is available for download at https://zenodo.org/record/3239543.
Barkopedia_Individual_Dog_Recognition_DatasetThis dataset is for Barkopedia Challenge https://uta-acl2.github.io/barkopedia.html
📦 Dataset Description
Check Training Data here: ArlingtonCL2/Barkopedia_Individual_Dog_Recognition_Dataset
This challenge provides a dataset of labeled dog bark audio clips:
8924
Training set: 7137 clips (~120 clips per 60 dog IDs).
Test set: 1787 clips (~30 clips per 60 dog IDs) with hidden labels:
709 public (~40%) for live leaderboard updates.
1078 private (~60%) for final evaluation.
You will… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/Barkopedia_Individual_Dog_Recognition_Dataset.DriveQA_Dataset
DriveQA: Passing the Driving Knowledge Test
Dataset Summary
DriveQA is a comprehensive multimodal benchmark that evaluates driving knowledge through text-based and vision-based question-answering tasks. The dataset simulates real-world driving knowledge tests, assessing LLMs and MLLMs on traffic regulations, sign recognition, and right-of-way reasoning.
Supported Tasks
Text-based QA: Traffic rules, safety regulations, right-of-way principles… See the full description on the dataset page: https://huggingface.co/datasets/barkmin77/DriveQA_Dataset.meeting-to-json-kopubchem-cid-smiles-title-inchikey-28M
Dataset Description
This dataset contains chemical molecular information in SMILES representation and other related metadata, extracted from the PubChem Compound Extras FTP directory.
Data Source
The SMILES data used to create this dataset can be found from the following PubChem FTP location:
PubChem Compound Extras
big-red-bark-chat-evaluation
Big Red Bark Chat Q&A Dataset
Dataset Description
This dataset contains 12,385 question-and-answer pairs collected from Big Red Bark Chat, an innovative AI assistant developed at Cornell University that answers questions about dog health (as well as other animal species). While it does not replace professional veterinary advice, it serves as a valuable starting point by searching trusted sources. Big Red Bark Chat is designed to provide quick and reliable answers… See the full description on the dataset page: https://huggingface.co/datasets/Sr523/big-red-bark-chat-evaluation.barkleyBarkopediaDogEmotionClassification_Data
🐶 Barkopedia Challenge Dataset
🔗 Barkopedia Website
📦 Dataset Description
This challenge provides a labeled dataset of dog bark audio clips for understanding the arousal and valence of emotional state from sound.
📁 Current Release
Training Set
Includes:
train/husky and train/shiba contain all training audio clips for each of the two breeds
husky_train_labels.csv and shiba_train_labels.csv contain labels for all training audio clips
Total:… See the full description on the dataset page: https://huggingface.co/datasets/codeet/BarkopediaDogEmotionClassification_Data.bark-spa-token-trainybark-detection
Bark detection dataset
Dataset Description
This dataset comprises both positive and negative samples of audio of 1 second in WAV format, recorded at 44.1kHz.
Negative samples include music, voice, claps, whistles and vacuum cleaner noise, among other sound you may record inside a house.
Caveats:
This is an imbalanced dataset: ~10k negatives vs ~500 positives.
Positive samples may include human generated barks.
Some (few) positive samples are false positives.… See the full description on the dataset page: https://huggingface.co/datasets/rmarcosg/bark-detection.details_Slim205__Barka-2b-it_v2_alrage
Dataset Card for Evaluation run of Slim205/Barka-2b-it
Dataset automatically created during the evaluation run of model Slim205/Barka-2b-it.
The dataset is composed of 1 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Slim205__Barka-2b-it_v2_alrage.SWIFTT-bark_beetle_detection_semantic_segmentation
Bark Beetle Detection - Semantic Segmentation Dataset (SWIFTT Project)
📋 Overview
This collection of datasets is devoleped in fullfilment of the research objectives of SWIFTT project (Satellites for Wilderness Inspection and Forest Threat Tracking),
funded by the European Union under Grant Agreement 101082732.
It contains labeled satellite imagery acquired with Sentinel-2 and Sentinel-1, that can be used for developing and evaluating predictive models to detect… See the full description on the dataset page: https://huggingface.co/datasets/AnnalisaAppice/SWIFTT-bark_beetle_detection_semantic_segmentation.one-piece-character-birthdays
