datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BeehiveStatesClassification
Beehive States for MTEB
This artifact contains the two official full-mode Beehive States folds from
HEAR release 2021.3. Each fold contains the same 576 ten-minute, mono, 48 kHz
PCM16 NU-Hive recordings, but preserves HEAR's released preprocessing output
and cross-hive split independently. Fold 0 trains and validates on Hive 1 and
tests on Hive 3; fold 1 reverses the hives. Labels are NOQUEEN and QUEEN.
The 576 IDs and labels agree across folds, but none of the corresponding WAV… See the full description on the dataset page: https://huggingface.co/datasets/artist/BeehiveStatesClassification.trace-bench
TRACE-Bench: Trustworthy Audio Cue Evaluation Benchmark
TRACE-Bench is a large-scale, controlled audio benchmark for evaluating the trustworthiness of Audio Language Models (ALMs) across four dimensions: Safety, Fairness, Robustness, and Privacy. It is, to our knowledge, the first benchmark to systematically attribute trustworthiness failures to specific acoustic cue types and ALM architectural components.
Overview
Property
Value
Total audio instances
37… See the full description on the dataset page: https://huggingface.co/datasets/BeelieverBzz/trace-bench.UZ_voicespeech_commands--> This is an exact copy of google/speech_commands adapted to be usable with recent datasets 🤗 versions (no remote code). <--
Dataset Card for SpeechCommands
Dataset Summary
This is a set of one-second .wav audio files, each containing a single spoken
English word or background noise. These words are from a small set of commands, and are spoken by a
variety of different speakers. This data set is designed to help train simple
machine learning models. It is covered in… See the full description on the dataset page: https://huggingface.co/datasets/beeneptune/speech_commands.uzbek_speech_dataSTT_uzThe dataset is organized into the following directories and files:
audio/
other/: Contains .tar archives like uz_other_0.taruz_other_1.tar
train/: Contains .tar archives like uz_train_0.tar.
validated/: Contains .tar archives like uz_validated_0.tar, uz_validated_1.tar, and uz_validated_2.tar.
test/: Contains individual .wav files.
transcription/: Contains .tsv files including:
other.tsv
train.tsv
validated.tsv
test.tsv
The .tsv files have two columns: file_name and transcription. Each entry… See the full description on the dataset page: https://huggingface.co/datasets/Beehzod/STT_uz.fongbe-speechuzbek_stt_databeepbank-500
BeepBank-500 (earcons-mini-500)
BeepBank-500 is a compact, fully synthetic earcon/alert mini‑dataset (≈300–500 clips) for UI sound research.
It contains short tones and triads generated from a controlled parameter grid (waveform family, f0, duration,
envelope, amplitude modulation, and simple Schroeder-style reverbs). The dataset ships with a metadata schema,
lightweight baselines, and a data note template for arXiv. Audio is intended for release under CC0-1.0 (public domain).
Code… See the full description on the dataset page: https://huggingface.co/datasets/mandipgoswami/beepbank-500.new_datasetHEARBeehiveStatesClassification_BeehiveStatesfongbe-speech-dataset-female
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: Fréjus LALEYE
Shared by [optional]: Fréjus LALEYE
Language(s) (NLP): Fongbe
License: [More Information Needed]
Dataset Sources [optional]
Repository: https://github.com/laleye/pyFongbe
Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/beethogedeon/fongbe-speech-dataset-female.tts_2301beethoven-grandstaff-multimodaldataset_for_STT_TTSmodelsbot_sttbee-islandUzTTS_datauz-datafongbe-speech-dataset-malebeekeeping_tech_hi
Dataset Card for "beekeeping_tech_hi"
More Information needed
beethoven-quartetsADA_uz_dataaug_uzbek_datasetbeethoven
Beethoven Sonatas Dataset
Beethoven is a raw audio waveform dataset used in the paper "It's Raw! Audio Generation with State-Space Models". It has been used primarily as a source of single instrument piano music for training music generation models at a small scale.
The dataset was originally introduced in the SampleRNN paper by Mehri et al. (2017) and download details from the original paper can be found at… See the full description on the dataset page: https://huggingface.co/datasets/krandiash/beethoven.dataset_TTS_uzbeethoven-reference-audioamandaaudio
