mandi
Datasets
All datasets matching “mandi”rirmega
RIRmega v2 — Dataset card (Hugging Face)
This folder holds the v2 artifacts intended for the Hugging Face dataset mandipgoswami/rirmega. When published, use revision v2.0.0 for the v2 release.
Dataset description
RIRmega v2 extends the existing RIRmega v1 dataset with:
A versioned metadata schema (metadata_v2.parquet) with acoustic metrics (RT60, DRR, C50, C80, D50, EDT), quality-control grades, and provenance.
A QC report (qc_report.parquet) with checks and outlier… See the full description on the dataset page: https://huggingface.co/datasets/mandipgoswami/rirmega.whisper-rirmega-bench
Whisper-RIR-Mega: Paired Clean↔Reverberant Speech Robustness Benchmark
Dataset Summary
Whisper-RIR-Mega is a benchmark dataset of paired clean and reverberant speech for evaluating ASR robustness to room acoustics. Each sample consists of:
audio_clean: Clean speech (LibriSpeech test-clean, 16 kHz)
audio_reverb: Same utterance convolved with one RIR from RIR-Mega (v2)
text_ref: Ground-truth transcript
RIR metadata: rir_id, RT60, DRR, C50, etc. when available
Technical… See the full description on the dataset page: https://huggingface.co/datasets/mandipgoswami/whisper-rirmega-bench.AnomalyMachine-50K
Dataset Summary
AnomalyMachine-50K is a fully synthetic industrial machine sound anomaly detection dataset designed for research on acoustic monitoring, predictive maintenance, and sound event detection.The dataset contains 50,000 monaural audio clips, each 10 seconds long at 22,050 Hz, covering six industrial machine types, multiple operating conditions, and diverse anomaly types under different signal-to-noise ratios.
The dataset is generated entirely via signal-processing based… See the full description on the dataset page: https://huggingface.co/datasets/mandipgoswami/AnomalyMachine-50K.VACE-Benchmark
VACE: All-in-One Video Creation and Editing
(ICCV 2025)
Zeyinzi Jiang*
·
Zhen Han*
·
Chaojie Mao*†
·
Jingfeng Zhang
·
Yulin Pan
·
Yu Liu
Tongyi Lab -
Introduction
VACE is an all-in-one model designed for video creation and editing. It encompasses various tasks, including reference-to-video generation (R2V), video-to-video editing (V2V), and masked video-to-video editing… See the full description on the dataset page: https://huggingface.co/datasets/Manding0/VACE-Benchmark.LibriRIR-100
LibriRIR-100
Dataset Summary
LibriRIR-100 is a large-scale paired clean↔reverberant speech training corpus containing exactly 100 hours of speech. Each utterance is paired with a room impulse response from RIR-Mega (mandipgoswami/rirmega), stratified across four RT60 reverberation conditions. Designed as a drop-in training resource for robust ASR, speech enhancement, and dereverberation models.
Why LibriRIR-100
Existing paired reverberant speech datasets are… See the full description on the dataset page: https://huggingface.co/datasets/mandipgoswami/LibriRIR-100.rir-mega-speech
RIR-Mega-Speech
Dataset Summary
RIR-Mega-Speech is a large-scale reverberant speech corpus created by convolving LibriSpeech utterances with simulated room impulse responses (RIRs sampled from the RIR-Mega collection). Each reverberant utterance includes per-file acoustic metadata computed from the source RIR, enabling controlled analysis of reverberation effects on speech processing systems.
This dataset emphasizes transparency and reproducibility: acoustic metrics are… See the full description on the dataset page: https://huggingface.co/datasets/mandipgoswami/rir-mega-speech.
