datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
prof_report__wavymulder-Analog-Diffusion__multi__24
Dataset Card for "prof_report__wavymulder-Analog-Diffusion__multi__24"
More Information needed
wavepulse-radio-raw-transcripts
WavePulse Radio Raw Transcripts
Dataset Summary
WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.extracel_waveforms
Spike waveform shards (derived from IBL dandiset 000409)
A curated collection of per-spike multichannel waveform windows extracted from NWB assets in the DANDI dandiset 000409 (International Brain Laboratory, IBL). This repository contains the extractor and uploader used to produce Parquet shards; this README documents the derived dataset, its provenance, format, and how to reproduce it.
Short description
Parquet shards of spike waveform windows (channels × timesteps).… See the full description on the dataset page: https://huggingface.co/datasets/rokaijano/extracel_waveforms.wav2vec2_common_voice_accents_3Urdu-ONYX-WAV-kanade-Annotated
Urdu-ONYX-WAV-real-Annotated
Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,217
Total Duration: 42.77 hours
Average Duration: 5.87 seconds
Duration Range: 0.65s - 122.23s
Average Phonemes: 18.5 per sample
Average Kanade Tokens: 151.1 per sample
Global Embedding Dimension: 128
New Columns
This dataset adds the following columns:
duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.everyayah-wav
everyayah-wav — Quranic recitation audio mirror
Full-mushaf Quranic recitation audio at 16 kHz mono 16-bit WAV, re-encoded
from everyayah.com for ML / ASR research.
This dataset is intentionally audio-only — no transcription text and no
alignment timings. The canonical Quranic text is widely available from
Tanzil and other public sources; pair this audio with
whatever text edition fits your use case.
Schema
Column
Type
Notes
audio
Audio(16000)
16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/dev-ahmedhany/everyayah-wav.mls10k_wavtokenizertransfer_1.2_wave2vec
Dataset Card for "transfer_1.2_wave2vec"
More Information needed
datasetgravitational-waves-strainWaveForcing_Stage2_ODE_pair
WaveForcing Stage 2 — Wan2.1-T2V-14B ODE Endpoint Pairs 2K
本数据集包含 2,176 对文本条件视频生成端点数据:2,048 对训练数据及 128 对留出验证数据。每对数据保存文本提示词、初始高斯噪声 z_ref,以及同一提示词和噪声经 Wan2.1-T2V-14B teacher 去噪得到的最终 latent y_ref,可用于视频扩散模型的成对蒸馏或回归研究。
这里的 “ODE pairs” 指 初始噪声到最终去噪 latent 的端点对。数据没有保存 50 步采样过程中的中间状态,不是完整 ODE 轨迹数据集。
数据划分
Split
数量
pair_index
文件名
Train
2,048
0–2047
000000.pt–002047.pt
Validation
128
2048–2175
002048.pt–002175.pt
总计
2,176
0–2175
2,176 个 .pt 文件
划分由… See the full description on the dataset page: https://huggingface.co/datasets/Osc7/WaveForcing_Stage2_ODE_pair.wav2vec2-ru-IIIN4_slice
30-Day Canonical Tav RGC N4 Lattice & Telemetry Slices
253 prospect-named windows · parti-wave/N4_slice
Resonance Graph Core (RGC) v0.2 · 49-core logistic substrate · paired data & session logs
This dataset contains 253 overlapping time windows cut from a continuous 30-day logistic-map characterization run of a 49-core Resonance Graph Core (RGC) on a Digilent Arty A7-100T FPGA, plus the matching windows from the companion session log.
Each data slice and each log slice carries… See the full description on the dataset page: https://huggingface.co/datasets/parti-wave/N4_slice.wave-propagation-1d
Wave1D-Propagation — StructBench canonical dataset
Download
One case, one file — fetch exactly what you need (pip install huggingface_hub):
from huggingface_hub import hf_hub_download, snapshot_download
# one case
path = hf_hub_download("StructBench/wave-propagation-1d",
filename="<case_id>.h5", repo_type="dataset")
# the full archive (resumable; cached under HF_HOME)
root = snapshot_download("StructBench/wave-propagation-1d"… See the full description on the dataset page: https://huggingface.co/datasets/StructBench/wave-propagation-1d.elliott-wave-market-data
Elliott Wave Market Data
Comprehensive OHLCV market data for training Elliott Wave pattern recognition neural networks.
Dataset Description
This dataset contains historical OHLCV (Open, High, Low, Close, Volume) data across multiple asset classes and timeframes, specifically curated for Elliott Wave analysis and pattern recognition.
Asset Classes
Stocks: S&P 500 components and international equities
Crypto: Top 50 cryptocurrencies by market cap
ETFs: Sector… See the full description on the dataset page: https://huggingface.co/datasets/usamaahmedsh/elliott-wave-market-data.hplt-greek-ge8-no-mt-clean60-wave4
HPLT Greek GE8 No-MT Clean60 Wave4
A standalone release of the filtered Greek HPLT slice used in the GlossAPI Greek pretraining corpus. It contains the full HPLT/ell_Grek_ge8_no_mt_clean60 source after the Wave4 re-cleaning and normalization pass.
Snapshot
Rows: 48728774
Data parquet files: 250
Source dataset value: HPLT/ell_Grek_ge8_no_mt_clean60
Quality bins: 8, 9, 10
MT/register filtering: applied before this release
Cleaner gate: greek_badness_score <= 60 before… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/hplt-greek-ge8-no-mt-clean60-wave4.elliott-wave-market-data-complete
Elliott Wave Market Data - Complete (Quality Validated) ✅
Production-ready, quality-validated OHLCV market data for training Elliott Wave pattern recognition neural networks.
🎯 Key Features
Quality Validated: Rigorous data quality checks applied
Complete Coverage: 1,403 unique instruments
Multi-Timeframe: 1h, 4h, 1d, 1wk data
22,546,189 Total Data Points
Dataset Statistics
Timeframe
Rows
Tickers
1h
9,267,368
1,358
4h
2,849,583
1,358
1d
8,648… See the full description on the dataset page: https://huggingface.co/datasets/usamaahmedsh/elliott-wave-market-data-complete.elliott-wave-market-data-extended
Elliott Wave Market Data - Extended Dataset
Additional comprehensive OHLCV market data for training Elliott Wave pattern recognition neural networks.
⚠️ This is a companion dataset - contains instruments NOT included in the primary dataset.
Dataset Description
This extended dataset contains historical OHLCV data for additional asset classes and instruments, specifically curated to complement the primary Elliott Wave dataset.
Asset Classes (Different from Primary… See the full description on the dataset page: https://huggingface.co/datasets/usamaahmedsh/elliott-wave-market-data-extended.wav2vec2-ru-IIwav2vec2-ru-IVreachy_wave_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "reachy2",
"total_episodes": 50,
"total_frames": 15688,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aliberts/reachy_wave_2.SiN-photonic-waveguide-loss-efficiency
💎 SiN Photonic Waveguide Loss & Efficiency Dataset
🔬 90,000 synthetic rows of silicon nitride (Si₃N₄) waveguide parameters linking geometry, fabrication, and operating conditions to loss and efficiency metrics, for regression modeling, simulation, and fine-tuning.
⚠️ Disclaimer: All rows are synthetically generated. Parameter ranges are informed by published SiN platform values, but no row is a foundry measurement. The data_source column is a schema field; every row in this… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/SiN-photonic-waveguide-loss-efficiency.MELD-processed-v3-wavlm
MELD Processed Multi-Modal Emotion Recognition Dataset
Processed dataset containing Prosody, Whisper acoustic encodings, DistilBERT text hidden states, and Ekman emotion labels.
so101_grapepickplacegigaspeech_xl_wavtokenizergravitational-wave-events
Gravitational Wave Events (GWOSC)
Credit: NASA/CXC/A. Hobart
Part of a dataset collection on Hugging Face.
Dataset description
All confirmed gravitational wave events from the Gravitational-Wave Open Science Center (GWOSC), covering LIGO, Virgo, and KAGRA observing runs.
Gravitational waves are ripples in spacetime generated by the acceleration of massive objects, predicted by Einstein's general theory of relativity in 1916 and first directly detected on… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/gravitational-wave-events.librispeech960-wavlm-large-km1000_asr
Dataset Card for "librispeech960-wavlm-large-km1000_asr"
More Information needed
wave-archive-USA-southwest
Dataset Card for wave-archive-USA-southwest
A dataset of NDBC/NOAA stdmet and spectral readings since 1991 for the US-Southwest Pacific coast.
Dataset Details
Dataset Description
This is historical data from the NOAA/NDBC consisting of the following types:
Standard Meteorological: "stdmet
Continuous Wind: "cwind"
Spectral Wave Density: "swden"
Spectral Wave Direction (α₁): "swdir"
Directional Spreading (R₁): "swr1"
The spectral data is aligned by timestamp… See the full description on the dataset page: https://huggingface.co/datasets/surfe-diem/wave-archive-USA-southwest.wavebender_dataset
_ _ __ _ _ ____ ____ ____ _ _ ____ ____ ____
( \/\/ ) /__\( \/ )( ___)( _ \( ___)( \( )( _ \( ___)( _ \
) ( /(__)\\ / )__) ) _ < )__) ) ( )(_) ))__) ) /
(__/\__)(__)(__)\/ (____)(____/(____)(_)\_)(____/(____)(_)\_)
OVERVIEW
UNDER DEVELOPMENT
This dataset was generated using the WAVEBENDER app by webXOS, located in the /generator/ folder of this repo. Download WAVE BENDER
to create your own similar datasets.… See the full description on the dataset page: https://huggingface.co/datasets/webxos/wavebender_dataset.youtube-cc-by-music📺 YouTube-CC-BY-Music 📺
YouTube-CC-BY-Music is a comprehensive collection of metadata for 316,000 music tracks shared on YouTube.
If you want the version of this dataset including prompt, see https://huggingface.co/datasets/WaveGenAI/youtube-cc-by-music_annoted.
Content
The dataset includes descriptions, tags, and other metadata associated with 316,000 music videos uploaded to YouTube under the CC-BY license. These videos come from a diverse range of artists and genres, providing… See the full description on the dataset page: https://huggingface.co/datasets/WaveGenAI/youtube-cc-by-music.
