fluid
Datasets
All datasets matching “fluid”fluidgym-datamusan
MUSAN: A Music, Speech, and Noise Corpus
MUSAN is a corpus of music, speech, and noise recordings designed for training models for voice activity detection and music/speech discrimination. This is a comprehensive collection suitable for various audio processing tasks.
Dataset Structure
The dataset is organized into three main categories:
1. Music (~42 hours)
Subcategories: Classical, Pop/Rock, Jazz, and more
Sources: Free Music Archive, Jamendo, and others… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/musan.fleurs-full
FLEURS Full - Test Set for ASR Benchmarking
Complete test set of Google FLEURS for all 30 languages supported by Qwen3-ASR, prepared for benchmarking with FluidAudio.
Languages (30)
Asian Languages (13)
Code
Language
Samples
cmn_hans_cn
Chinese (Mandarin)
945
yue_hant_hk
Cantonese
819
ja_jp
Japanese
650
ko_kr
Korean
382
vi_vn
Vietnamese
857
th_th
Thai
1,021
id_id
Indonesian
687
ms_my
Malay
749
hi_in
Hindi
418
ar_eg
Arabic (Egyptian)… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/fleurs-full.ami-corpus-mirror
AMI Corpus Mirror
Mirror of the subset of the AMI Meeting Corpus used by
FluidAudio diarization benchmarks. Hosted here so CI and
local benchmark runs do not depend on the availability of the upstream groups.inf.ed.ac.uk server
(see FluidAudio#752).
Contents
annotations/ami_public_manual_1.6.2.zip — AMI public manual annotations v1.6.2
(repackaged from the official archive; identical content, including segments/, words/,
corpusResources/meetings.xml)… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/ami-corpus-mirror.THCHS-30-tests
THCHS-30 Test Set
THCHS-30 test split for Mandarin Chinese speech recognition benchmarking.
Dataset Info
Language: Mandarin Chinese (zh-CN)
Samples: 2,495
Speakers: 10
Sample Rate: 16 kHz
License: Apache 2.0
Usage
from datasets import load_dataset
# After uploading to HuggingFace
dataset = load_dataset("your-username/thchs30-test")
# Example
print(dataset['train'][0])
# {
# 'audio': {'array': [...], 'sampling_rate': 16000, 'path': 'audio/D11_750.wav'},
#… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/THCHS-30-tests.HUGGER-Unified-Gravity-Fluid-Framework
🌍 H.U.G.G.E.R: Heuristic Universal Grid & Gravity Equilibrium Rendering Tensor
This repository serves as an open academic archive and tensor-specification benchmark for generalized tensor standards, designed to resolve non-linear computational collapse and topological pole singularities in high-performance CFD and planetary atmospheric models.
It acts as the Macroscopic Gravitational Backbone, perfectly entangled with the microscopic Topological Zero Tensor (TZT)… See the full description on the dataset page: https://huggingface.co/datasets/jskresearch/HUGGER-Unified-Gravity-Fluid-Framework.
