CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01safe-autonomous-systems /fluidgym-data0 likes5.5k downloads6mo agoHugging Face02FluidInference /musan MUSAN: A Music, Speech, and Noise Corpus MUSAN is a corpus of music, speech, and noise recordings designed for training models for voice activity detection and music/speech discrimination. This is a comprehensive collection suitable for various audio processing tasks. Dataset Structure The dataset is organized into three main categories: 1. Music (~42 hours) Subcategories: Classical, Pop/Rock, Jazz, and more Sources: Free Music Archive, Jamendo, and others… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/musan.audion<1K3 likes1.5k downloads1y agoHugging Face03FluidInference /fleurs-full FLEURS Full - Test Set for ASR Benchmarking Complete test set of Google FLEURS for all 30 languages supported by Qwen3-ASR, prepared for benchmarking with FluidAudio. Languages (30) Asian Languages (13) Code Language Samples cmn_hans_cn Chinese (Mandarin) 945 yue_hant_hk Cantonese 819 ja_jp Japanese 650 ko_kr Korean 382 vi_vn Vietnamese 857 th_th Thai 1,021 id_id Indonesian 687 ms_my Malay 749 hi_in Hindi 418 ar_eg Arabic (Egyptian)… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/fleurs-full.audio10K<n<100K0 likes1k downloads4mo agoHugging Face04FluidInference /ami-corpus-mirror AMI Corpus Mirror Mirror of the subset of the AMI Meeting Corpus used by FluidAudio diarization benchmarks. Hosted here so CI and local benchmark runs do not depend on the availability of the upstream groups.inf.ed.ac.uk server (see FluidAudio#752). Contents annotations/ami_public_manual_1.6.2.zip — AMI public manual annotations v1.6.2 (repackaged from the official archive; identical content, including segments/, words/, corpusResources/meetings.xml)… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/ami-corpus-mirror.audio1K<n<10K0 likes626 downloads3mo agoHugging Face05FluidInference /THCHS-30-tests THCHS-30 Test Set THCHS-30 test split for Mandarin Chinese speech recognition benchmarking. Dataset Info Language: Mandarin Chinese (zh-CN) Samples: 2,495 Speakers: 10 Sample Rate: 16 kHz License: Apache 2.0 Usage from datasets import load_dataset # After uploading to HuggingFace dataset = load_dataset("your-username/thchs30-test") # Example print(dataset['train'][0]) # { # 'audio': {'array': [...], 'sampling_rate': 16000, 'path': 'audio/D11_750.wav'}, #… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/THCHS-30-tests.audio1K<n<10K0 likes571 downloads6mo agoHugging Face06jskresearch /HUGGER-Unified-Gravity-Fluid-Framework 🌍 H.U.G.G.E.R: Heuristic Universal Grid & Gravity Equilibrium Rendering Tensor This repository serves as an open academic archive and tensor-specification benchmark for generalized tensor standards, designed to resolve non-linear computational collapse and topological pole singularities in high-performance CFD and planetary atmospheric models. It acts as the Macroscopic Gravitational Backbone, perfectly entangled with the microscopic Topological Zero Tensor (TZT)… See the full description on the dataset page: https://huggingface.co/datasets/jskresearch/HUGGER-Unified-Gravity-Fluid-Framework.documentroboticsn<1K1 likes501 downloads5d agoHugging Face07terrant20 /FluidNexusDatasets FluidNexus: 3D Fluid Reconstruction and Prediction From a Single Video Yue Gao*, Hong-Xing "Koven" Yu*, Bo Zhu, Jiajun Wu Stanford University; Microsoft; Georgia Institute of Technology * denotes equal contribution FluidNexus-Smoke and FluidNexus-Ball Our FluidNexus-Smoke and FluidNexus-Ball datasets each include 120 scenes. Every scene contains 5 synchronized multi-view videos, with cameras arranged along a horizontal arc of approximately 120°. FluidNexusSmoke… See the full description on the dataset page: https://huggingface.co/datasets/terrant20/FluidNexusDatasets.0 likes485 downloads8mo agoHugging Face08FluidVerse /samples Fluidverse Samples This dataset provides sample files for all of the datasets published by the Fluidverse Organization. For each dataset, two trajectories are provided - one with a single bubble / droplet and one multi-droplet / bubble trajectory with five bubbles / droplets. Note: This is not a standalone dataset with a test and train split. This huggingface dataset serves solely as a preview for the complete datasets listed below. Dataset samples Dataset… See the full description on the dataset page: https://huggingface.co/datasets/FluidVerse/samples.n<1K0 likes404 downloads5mo agoHugging Face09cmudrc /OpenSeeSimE-Fluid OpenSeeSimE-Fluid: Engineering Simulation Visual Question Answering Benchmark Dataset Summary OpenSeeSimE-Fluid is a large-scale benchmark dataset for evaluating vision-language models on computational fluid dynamics (CFD) simulation interpretation tasks. It contains approximately 98,000 question-answer pairs across parametrically-varied fluid simulations including turbulent flow, heat transfer, and complex flow patterns. Purpose While vision-language models… See the full description on the dataset page: https://huggingface.co/datasets/cmudrc/OpenSeeSimE-Fluid.imagevisual-question-answering10K<n<100K0 likes349 downloads9mo agoHugging Face10FluidInference /JSUT-basic5000 JSUT (Japanese Speech Corpus) - Test Subset A test subset of the JSUT corpus containing 500 Japanese utterances from the basic5000 dataset (BASIC5000_4501-5000). Dataset Structure jsut_ver1.1/ └── basic5000/ ├── wav/ # WAV audio files (500 files, 48kHz) ├── transcript_utf8.txt # Transcriptions └── recording_info.txt # Recording dates File Formats transcript_utf8.txt BASIC5000_4501:だが、エーアイセンター稼動を快く思わない...… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/JSUT-basic5000.audio1K<n<10K1 likes346 downloads6mo agoHugging Face11safe-autonomous-systems /fluidgym-experiments FluidGym Experiments Paper | GitHub | Documentation FluidGym is a standalone, fully differentiable benchmark suite for reinforcement learning (RL) in active flow control (AFC). Built entirely in PyTorch on top of the GPU-accelerated PICT solver, it provides standardized evaluation protocols and diverse environments for systematic comparison of control methods. This repository contains the training and test datasets with results for all experimental runs presented in the paper.… See the full description on the dataset page: https://huggingface.co/datasets/safe-autonomous-systems/fluidgym-experiments.reinforcement-learning2 likes302 downloads4mo agoHugging Face12FluidInference /librispeechaudio1K<n<10K0 likes271 downloads11mo agoHugging Face13fluid-concepts /sample-page-assets Sample-page assets Files the cards of the Fluid Concepts sample datasets on the Hub (ToolTalk, Multimodal Expert Instruction, Multimodal Peer Collaboration) need to show to visitors who have not requested access yet, which the gated sample repositories cannot serve themselves: the card banners (*.png); a public copy of each sample repository's TECHNICAL.md and TERMS.md, under the repository's name, refreshed on every push of that repository. Nothing else is published here. The… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/sample-page-assets.0 likes238 downloads3d agoHugging Face14cmudrc /OpenSeeSimE-Fluid-Small OpenSeeSimE-Fluid-Small A stratified 10% subset of cmudrc/OpenSeeSimE-Fluid for evaluating vision-language models at a reduced compute footprint while preserving the joint distribution of simulation type, question type, media type, and question id. Subset Provenance Parent dataset: cmudrc/OpenSeeSimE-Fluid (98,326 rows total) Rows in this subset: 9,881 (10.05% of parent) Source classes: Bent Pipe, Converging Nozzle, Heat Exchanger, Heat Sink, Mixing Pipe Parquet shards:… See the full description on the dataset page: https://huggingface.co/datasets/cmudrc/OpenSeeSimE-Fluid-Small.imagevisual-question-answering1K<n<10K0 likes220 downloads5mo agoHugging Face15FluidInference /fleurs FLEURS Test Dataset Reorganized FLEURS test dataset with audio and transcripts together. Structure fleurs-test/ ├── en_us/ │ ├── en_us_0000.wav │ ├── en_us_0001.wav │ ├── ... │ ├── en_us.trans.txt (LibriSpeech format) │ ├── en_us.csv (detailed metadata) │ └── en_us.json (JSON metadata) ├── fr_fr/ │ └── ... └── ... Languages bg_bg: 350 test samples cs_cz: 350 test samples da_dk: 930 test samples de_de: 350 test samples el_gr: 650… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/fleurs.audio2 likes212 downloads1y agoHugging Face16KrishMalik /deltav-fluid1 likes184 downloads10d agoHugging Face17fluid-concepts /tooltalk-samplesgated ToolTalk Samples - High Quality Duplex Speech and Tool-calling in Customer Service Domain Two people improvise realistic customer-service calls while one operates a live, stateful tool environment—with synchronized speaker-separated audio, tool calls, and outcomes. ▶ Listen to Clean · ▶ Listen to Noisy · Discuss the full dataset In this sample: 26 calls · 90.7 minutes · 7 sample domains · 207 tool calls Technical specs: 48 kHz / 32-bit PCM speaker-separated source… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/tooltalk-samples.audion<1K1 likes169 downloads3d agoHugging Face18yuegao /FluidNexusDatasets FluidNexus: 3D Fluid Reconstruction and Prediction From a Single Video Yue Gao*, Hong-Xing "Koven" Yu*, Bo Zhu, Jiajun Wu Stanford University; Microsoft; Georgia Institute of Technology * denotes equal contribution FluidNexus-Smoke and FluidNexus-Ball Our FluidNexus-Smoke and FluidNexus-Ball datasets each include 120 scenes. Every scene contains 5 synchronized multi-view videos, with cameras arranged along a horizontal arc of approximately 120°. FluidNexusSmoke… See the full description on the dataset page: https://huggingface.co/datasets/yuegao/FluidNexusDatasets.2 likes166 downloads1y agoHugging Face19FluidInference /cv-corpus-25.0-ja Mozilla Common Voice 25.0 - Japanese Test Set (Complete) Dataset Description Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace. Key Features Size: 9,019 validated test utterances Coverage: 100% of official Common Voice 25.0 Japanese test split Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.audioautomatic-speech-recognition1K<n<10K0 likes162 downloads6mo agoHugging Face20allenai /fluid-benchmarking Fluid Language Model Benchmarking This dataset provides IRT models for ARC Challenge, GSM8K, HellaSwag, MMLU, TruthfulQA, and WinoGrande. Furthermore, it contains results for pretraining checkpoints of Amber-6.7B, K2-65B, OLMo1-7B, OLMo2-7B, Pythia-2.8B, and Pythia-6.9B, evaluated on these six benchmarks. 🚀 Usage For utilities to use the dataset and to replicate the results from the paper, please see the corresponding GitHub… See the full description on the dataset page: https://huggingface.co/datasets/allenai/fluid-benchmarking.3 likes147 downloads1y agoHugging Face21fluid-concepts /friend-bench Can a model — or a human — tell how two people are related from a 20-second clip of how they interact? 🌐 Built on Seamless Interaction FriendBench is a suite of benchmarks for social perception from thin-slice dyadic interaction — inferring facts about two people's relationship from a brief clip of how they interact, built on the Seamless Interaction dataset. Each released set is a config of this repository. 🎧 Multi-modal — text, audio, and video for every clip 🎯 Objective label —… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/friend-bench.audioaudio-classificationn<1K1 likes145 downloads2mo agoHugging Face22OzTianlu /Reasoning_as_Fluid Reasoning as Fluid: The Minimal Primitives and Inevitable Collapse of Linear Space Representation Author: Zixi "Oz" Li (Independent Researcher) Publication Date: November 2025 DOI: 10.57967/hf/7081 arXiv: (Submitted) PDF: fluid_reasoning.pdf 📋 Abstract We establish that reasoning is fundamentally a fluid-like dynamical process on constrained manifolds, not a computation in linear vector spaces. Through rigorous mathematical proof, we demonstrate: Minimal Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/OzTianlu/Reasoning_as_Fluid.documentn<1K1 likes117 downloads10mo agoHugging Face23fluid-concepts /multimodal-peer-collaboration-samplesgated Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges. ▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.audion<1K1 likes114 downloads3d agoHugging Face24FluidVerse /2D_SABW_OOOO 2D Shock-Induced Air Bubble Collapse in Water with Open BC Description The dataset captures the time-evolving behavior of 2D cylindrical air bubbles subjected to an external shock wave in water. The interaction with the shock wave results in a collapse of the air bubble. Here we investigate a scenario with open boundary conditions at every wall. About the data Metadata Description Solver ALPACA PDE 2D compressible Euler equations… See the full description on the dataset page: https://huggingface.co/datasets/FluidVerse/2D_SABW_OOOO.video10K<n<100K0 likes109 downloads5mo agoHugging Face25fluid-concepts /multimodal-expert-instruction-samplesgated Multimodal Expert Instruction Samples - Musical Instrument Lessons with Channel-separated Audio and Video A music teacher and a student work through two one-on-one lessons: both voices and both instruments on separate tracks, the student on camera, with the lesson plans, the instructions given to each side and both sides' post-lesson ratings alongside. ▶ Watch the lessons · See Peer Collaboration samples · Discuss the full collection Sister collection: Peer Collaboration… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-expert-instruction-samples.audion<1K1 likes98 downloads3d agoHugging Face26FluidVerse /2D_SDBA_SSOO 2D Shock-Induced Droplet Breakup in Air with Symmetric BC Description The dataset captures the time-evolving behavior of 2D cylindrical droplets subjected to an external shock wave in air. The interaction with the shock wave results in two different breakup modes namely SIE and RTP depending on the weber number. Here we investigate a scenario with symmetric boundary conditions at the north and south walls. About the data Metadata Description… See the full description on the dataset page: https://huggingface.co/datasets/FluidVerse/2D_SDBA_SSOO.video10K<n<100K0 likes91 downloads5mo agoHugging Face27FluidVerse /2D_SRBA_OOOO 2D Shock-Induced R22 Bubble Collapse in Air with Open BC Description The dataset captures the time-evolving behavior of 2D cylindrical R22 bubbles subjected to an external shock wave in air. The interaction with the shock wave results in a collapse of the R22 bubble. Here we investigate a scenario with open boundary conditions at every wall. About the data Metadata Description Solver ALPACA PDE 2D compressible Euler equations Dimension… See the full description on the dataset page: https://huggingface.co/datasets/FluidVerse/2D_SRBA_OOOO.10K<n<100K0 likes64 downloads5mo agoHugging Face28SURF-FluidSimulation /FluidSimulation SURF: A Generalisation Benchmark for GNNs Predicting Fluid Dynamics SURF, is a benchmark designed to test the generalization of learned graph-based fluid simulators. The benchmark consists of seven independent datasets: Base Turned Topo Range Dynamic Full FullFiner Each dataset is available as separate *.zip file and consists of at least 1200 2D incompressible fluid flow simulations with 300 timesteps. The data structure is as follows: folder: dataset_name folders: dpx files:… See the full description on the dataset page: https://huggingface.co/datasets/SURF-FluidSimulation/FluidSimulation.2 likes54 downloads3y agoHugging Face29AIM-Intelligence /fluid-reasoning-representation-phase1 Fluid Reasoning Representation - Phase 1 Multi-Model + Cross-Domain Sweep Phase 1 artifacts for the ARR 2026 rebuttal of Fluid Reasoning Representation (Hook et al.). This dataset extends the original QwQ x Mystery Blocksworld study with: Second large reasoning model: Llama-3.3-Nemotron-Super-49B-v1 Two new domains: Mystery Logistics (PDDL Logistics with obfuscated action / predicate vocabulary) and GSM8K-Renamed (math word problems with surface noun + verb obfuscation). C3 causal… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/fluid-reasoning-representation-phase1.text-generation0 likes52 downloads4mo agoHugging Face30cmudrc /OpenSeeSimE-Fluid-Mini OpenSeeSimE-Fluid-Mini A stratified 1% subset of cmudrc/OpenSeeSimE-Fluid for evaluating vision-language models at a reduced compute footprint while preserving the joint distribution of simulation type, question type, media type, and question id. Subset Provenance Parent dataset: cmudrc/OpenSeeSimE-Fluid (98,326 rows total) Rows in this subset: 1,040 (1.06% of parent) Source classes: Bent Pipe, Converging Nozzle, Heat Exchanger, Heat Sink, Mixing Pipe Parquet shards: 3… See the full description on the dataset page: https://huggingface.co/datasets/cmudrc/OpenSeeSimE-Fluid-Mini.imagevisual-question-answering1K<n<10K0 likes51 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.