datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WavCaps
WavCaps
WavCaps is a ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research, where the audio clips are sourced from three websites (FreeSound, BBC Sound Effects, and SoundBible) and a sound event detection dataset (AudioSet Strongly-labelled Subset).
Paper: https://arxiv.org/abs/2303.17395
Github: https://github.com/XinhaoMei/WavCaps
Statistics
Data Source
# audio
avg. audio duration (s)avg. text length
FreeSound… See the full description on the dataset page: https://huggingface.co/datasets/cvssp/WavCaps.wavelet-lstm-camels-models
Wavelet-LSTM CAMELS Streamflow Models
A collection of 61,380 pre-trained LSTM models for daily streamflow forecasting across 620 USGS catchments from the CAMELS dataset.
Each catchment has 99 independently trained models:
33 wavelet filters × 3 lead times (1, 3, 5 days) = 99 wavelet-enhanced models
33 matching baseline models (same architecture, no wavelet transform)
Models are designed to be ensembled across wavelets for robust predictions with uncertainty estimates.… See the full description on the dataset page: https://huggingface.co/datasets/johnswyou/wavelet-lstm-camels-models.prof_report__wavymulder-Analog-Diffusion__multi__24
Dataset Card for "prof_report__wavymulder-Analog-Diffusion__multi__24"
More Information needed
2ca36481WaveUI-25k
Dataset Card for WaveUI-25k
This is a FiftyOne dataset with 24977 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/WaveUI-25k")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/WaveUI-25k.wavefake-audiowavepulse-radio-raw-transcripts
WavePulse Radio Raw Transcripts
Dataset Summary
WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.extracel_waveforms
Spike waveform shards (derived from IBL dandiset 000409)
A curated collection of per-spike multichannel waveform windows extracted from NWB assets in the DANDI dandiset 000409 (International Brain Laboratory, IBL). This repository contains the extractor and uploader used to produce Parquet shards; this README documents the derived dataset, its provenance, format, and how to reproduce it.
Short description
Parquet shards of spike waveform windows (channels × timesteps).… See the full description on the dataset page: https://huggingface.co/datasets/rokaijano/extracel_waveforms.wav2vec2_common_voice_accents_33f9df09aWave-Anomaly-Detectionwavepulse-radio-summarized-transcripts
WavePulse Radio Summarized Transcripts
Dataset Summary
WavePulse Radio Summarized Transcripts is a large-scale dataset containing summarized transcripts from 396 radio stations across the United States, collected between June 26, 2024, and October 3, 2024. The dataset comprises approximately 1.5 million summaries derived from 485,090 hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The raw version of the transcripts is available… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-summarized-transcripts.Urdu-ONYX-WAV-kanade-Annotated
Urdu-ONYX-WAV-real-Annotated
Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,217
Total Duration: 42.77 hours
Average Duration: 5.87 seconds
Duration Range: 0.65s - 122.23s
Average Phonemes: 18.5 per sample
Average Kanade Tokens: 151.1 per sample
Global Embedding Dimension: 128
New Columns
This dataset adds the following columns:
duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.everyayah-wav
everyayah-wav — Quranic recitation audio mirror
Full-mushaf Quranic recitation audio at 16 kHz mono 16-bit WAV, re-encoded
from everyayah.com for ML / ASR research.
This dataset is intentionally audio-only — no transcription text and no
alignment timings. The canonical Quranic text is widely available from
Tanzil and other public sources; pair this audio with
whatever text edition fits your use case.
Schema
Column
Type
Notes
audio
Audio(16000)
16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/dev-ahmedhany/everyayah-wav.odinAdapted from ODIN (the Online Database of INterlinear glossed text). Adapted to the SIGMORPHON-2023 interlinear gloss shared task format by Nate Robinson.
Citations
Adapted Corpus
@inproceedings{he-etal-2023-sigmorefun,
title = "{S}ig{M}ore{F}un Submission to the {SIGMORPHON} Shared Task on Interlinear Glossing",
author = "He, Taiqi and
Tjuatja, Lindia and
Robinson, Nathaniel and
Watanabe, Shinji and
Mortensen, David R. and… See the full description on the dataset page: https://huggingface.co/datasets/wav2gloss/odin.bridge_module_wav2vec_librispeech_v1wave-uiLICENSE
2489f877wave-ui-25k
WaveUI-25k
This dataset contains 25k examples of labeled UI elements. It is a subset of a collection of ~80k preprocessed examples assembled from the following sources:
WebUI
RoboFlow
GroundUI-18K
These datasets were preprocessed to have matching schemas and to filter out unwanted examples, such as duplicated, overlapping and low-quality datapoints. We also filtered out many text elements which were not in the main scope of this work.
The WaveUI-25k dataset includes the original… See the full description on the dataset page: https://huggingface.co/datasets/agentsea/wave-ui-25k.WavDatasetwavenet_flashback
Dataset Card for "wavenet_flashback"
https://cloud.google.com/text-to-speech/docs/reference/rest/v1/text/synthesize#AudioConfig
sv-SE-Wavenet-{voice}
https://spraakbanken.gu.se/resurser/flashback-dator
wavcaps-10s-16k
wavcaps shorter than 10s & resample to 16k
sada-train-wav2vec2-xls-r-300m-ar-preprocessedyoutube_wavwavlm-large_layer21-librispeech-asr100h
Dataset Card for "wavlm-large_layer21-librispeech-asr100h"
More Information needed
DeepSeek-V4-Flash-0731-REAM-calibration-stats
DeepSeek-V4-Flash-0731 — expert calibration statistics (REAM line)
Layerwise routed-expert statistics of
deepseek-ai/DeepSeek-V4-Flash-0731
(43 MoE layers × 256 experts), collected by running the full model over a
~4.9M-token multi-domain calibration mix (multi-turn dialogs, thinking and
direct modes, rendered with the model's own chat encoder). These are the
statistics behind the REAM144/96 release line — published so that expert
selection, pruning, merging and routing research… See the full description on the dataset page: https://huggingface.co/datasets/WaveCut/DeepSeek-V4-Flash-0731-REAM-calibration-stats.MusicCaps_30s_wavlvdec-flf-wave-20260815-generated
LVDEC FLF Wave 2026-08-15 — Generated Returns
This public dataset contains the completed generated return package for the
Yueshengli/lvdec-flf-wave-20260815
execution wave.
Contents
The dataset contains 8,193 completed jobs with no final failures:
Generator / batch
Jobs
ltx23_flf2v
1,200
ltx23_flf2va
1,200
ltx25_flf2v
470
ltx25_flf2va
370
wan22_fun_5b_control
990
minimax_h3_flf2v
1,166
minimax_h3_flf2va
2,797
Each job directory may… See the full description on the dataset page: https://huggingface.co/datasets/Hallucinatie/lvdec-flf-wave-20260815-generated.mls10k_wavtokenizerver_waves
