datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ami-corpus-mirror
AMI Corpus Mirror
Mirror of the subset of the AMI Meeting Corpus used by
FluidAudio diarization benchmarks. Hosted here so CI and
local benchmark runs do not depend on the availability of the upstream groups.inf.ed.ac.uk server
(see FluidAudio#752).
Contents
annotations/ami_public_manual_1.6.2.zip — AMI public manual annotations v1.6.2
(repackaged from the official archive; identical content, including segments/, words/,
corpusResources/meetings.xml)… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/ami-corpus-mirror.THCHS-30-tests
THCHS-30 Test Set
THCHS-30 test split for Mandarin Chinese speech recognition benchmarking.
Dataset Info
Language: Mandarin Chinese (zh-CN)
Samples: 2,495
Speakers: 10
Sample Rate: 16 kHz
License: Apache 2.0
Usage
from datasets import load_dataset
# After uploading to HuggingFace
dataset = load_dataset("your-username/thchs30-test")
# Example
print(dataset['train'][0])
# {
# 'audio': {'array': [...], 'sampling_rate': 16000, 'path': 'audio/D11_750.wav'},
#… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/THCHS-30-tests.OpenSeeSimE-Fluid
OpenSeeSimE-Fluid: Engineering Simulation Visual Question Answering Benchmark
Dataset Summary
OpenSeeSimE-Fluid is a large-scale benchmark dataset for evaluating vision-language models on computational fluid dynamics (CFD) simulation interpretation tasks. It contains approximately 98,000 question-answer pairs across parametrically-varied fluid simulations including turbulent flow, heat transfer, and complex flow patterns.
Purpose
While vision-language models… See the full description on the dataset page: https://huggingface.co/datasets/cmudrc/OpenSeeSimE-Fluid.librispeechOpenSeeSimE-Fluid-Small
OpenSeeSimE-Fluid-Small
A stratified 10% subset of cmudrc/OpenSeeSimE-Fluid for evaluating vision-language models at a reduced compute footprint while preserving the joint distribution of simulation type, question type, media type, and question id.
Subset Provenance
Parent dataset: cmudrc/OpenSeeSimE-Fluid (98,326 rows total)
Rows in this subset: 9,881 (10.05% of parent)
Source classes: Bent Pipe, Converging Nozzle, Heat Exchanger, Heat Sink, Mixing Pipe
Parquet shards:… See the full description on the dataset page: https://huggingface.co/datasets/cmudrc/OpenSeeSimE-Fluid-Small.tooltalk-samples
ToolTalk Samples - High Quality Duplex Speech and Tool-calling in Customer Service Domain
Two people improvise realistic customer-service calls while one operates a live, stateful tool environment—with synchronized speaker-separated audio, tool calls, and outcomes.
▶ Listen to Clean · ▶ Listen to Noisy · Discuss the full dataset
In this sample: 26 calls · 90.7 minutes · 7 sample domains · 207 tool calls
Technical specs: 48 kHz / 32-bit PCM speaker-separated source… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/tooltalk-samples.cv-corpus-25.0-ja
Mozilla Common Voice 25.0 - Japanese Test Set (Complete)
Dataset Description
Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace.
Key Features
Size: 9,019 validated test utterances
Coverage: 100% of official Common Voice 25.0 Japanese test split
Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.friend-bench
Can a model — or a human — tell how two people are related from a 20-second clip of how they interact?
🌐 Built on Seamless Interaction
FriendBench is a suite of benchmarks for social perception from thin-slice dyadic
interaction — inferring facts about two people's relationship from a brief clip of how they
interact, built on the Seamless Interaction
dataset. Each released set is a config of this repository.
🎧 Multi-modal — text, audio, and video for every clip
🎯 Objective label —… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/friend-bench.multimodal-peer-collaboration-samples
Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles
Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges.
▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection
Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.multimodal-expert-instruction-samples
Multimodal Expert Instruction Samples - Musical Instrument Lessons with Channel-separated Audio and Video
A music teacher and a student work through two one-on-one lessons: both voices and both instruments on separate tracks, the student on camera, with the lesson plans, the instructions given to each side and both sides' post-lesson ratings alongside.
▶ Watch the lessons · See Peer Collaboration samples · Discuss the full collection
Sister collection: Peer Collaboration… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-expert-instruction-samples.OpenSeeSimE-Fluid-Mini
OpenSeeSimE-Fluid-Mini
A stratified 1% subset of cmudrc/OpenSeeSimE-Fluid for evaluating vision-language models at a reduced compute footprint while preserving the joint distribution of simulation type, question type, media type, and question id.
Subset Provenance
Parent dataset: cmudrc/OpenSeeSimE-Fluid (98,326 rows total)
Rows in this subset: 1,040 (1.06% of parent)
Source classes: Bent Pipe, Converging Nozzle, Heat Exchanger, Heat Sink, Mixing Pipe
Parquet shards: 3… See the full description on the dataset page: https://huggingface.co/datasets/cmudrc/OpenSeeSimE-Fluid-Mini.fluid-knowledge-validation
Fluid Knowledge
Public synthetic release-validation fixtures. These repeated arithmetic items test artifact publication and verification only; they were not authored or blindly reviewed by frontier models and are not a usable benchmark.
Each immutable epochs/<id>/manifest.json binds its published artifacts. commitment.json reveals the nonce for verification. Protocol and source attribution accompany each epoch. Pin the returned Hugging Face commit SHA for reproduction. Public… See the full description on the dataset page: https://huggingface.co/datasets/actuallymentor/fluid-knowledge-validation.clinical-fluid-balance-renal-response-coherence-risk-v0.1What this repo is for
Detect when
fluid balance signals
and
renal management
fall out of alignment
before
acute kidney injury
or fluid overload harm.
Testclinical-fluid-balance-order-monitoring-coherence-risk-v0.1What this repo is for
Detect when
fluid monitoring is ordered
but charting
is missing
incomplete
or too infrequent
before
AKI risk
fluid overload
missed deterioration
fluid-2-sft-eval
Fluid 2 — dictation cleanup eval
This repository contains 7,161 text-only voice-dictation cleanup evaluation
rows for like-for-like model comparison. No audio is included or fetched.
Split
Rows
Documents
Audio
eval
7,161
2,718
Not included
What this benchmark tests
The benchmark measures whether a model can turn noisy voice dictation into the
intended written text without answering it or adding content. The rows cover:
local ASR, spelling… See the full description on the dataset page: https://huggingface.co/datasets/johnbean393/fluid-2-sft-eval.Fluid-Mechanics-CoT
🌊 Engineering Fluid Mechanics CoT Dataset (工程流体力学思维链数据集)
📖 Dataset Description (数据集简介)
This dataset focuses on Engineering Fluid Mechanics, specifically designed to enhance Large Language Models' (LLMs) reasoning capabilities in complex physics problems.
Unlike standard QA datasets, this dataset provides Chain-of-Thought (CoT) annotations, breaking down the problem-solving process into:
Analysis & Reasoning: Strategy selection and physical law identification.… See the full description on the dataset page: https://huggingface.co/datasets/zshiyi/Fluid-Mechanics-CoT.clinical-fluid-balance-kidney-function-coherence-risk-v0.1What this repo is for
Detect when
fluid management
and kidney response
decouple
Common breaks
no fluid balance recorded in AKI risk
rising creatinine with no plan
plan documented but not executed
diuresis continues despite kidney deterioration
Used for
AKI safety
ICU and ward escalation
fluid governance
clinical-fluid-order-input-output-coherence-risk-v0.1What this repo is for
Detect when
IV fluids are prescribed or running
but I/O, weight, and balance review
do not support safe control
Common breaks
fluids running with no input chart
no output chart
no daily weight in risk patients
balance recorded but no plan change
order placed but not running
Examples
sepsis fluids continue with no urine chart
elderly patient gains weight but no review
renal patient on fluids with no net balance note
You use it to flag
fluid overload risk
AKI risk… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-fluid-order-input-output-coherence-risk-v0.1.clinical-fluid-balance-renal-function-coherence-risk-v0.1What this repo is for
Detect when fluid strategy
and renal function signals
fall out of alignment
before
fluid overload
or avoidable AKI.
clinical-fluid-balance-order-output-coherence-risk-v0.1What this repo is for
Detect when
fluid orders
and
actual patient output
fall out of alignment
before
fluid overload
or organ injury.
OpenSeeSimE-Fluid-Testfluid-2-sft-asr
Fluid 2 — synthetic dictation cleanup
Fluid 2 is an English supervised-fine-tuning corpus for models that turn noisy automatic-speech-recognition output into the written insertion a user intended. It contains 354,549 rows in official document-grouped 96/2/2 splits, 861.3 hours of processed 16 kHz speech, and 8.48M target-side loss tokens in 355 Parquet shards (49.25 GiB).
This is not an ordinary transcription dataset. The model sees document context plus an ASR hypothesis and… See the full description on the dataset page: https://huggingface.co/datasets/johnbean393/fluid-2-sft-asr.odia-critical-care-trauma-fluid-maths-reasoning-v3FluidFlowMLProcessedclinical-fluid-balance-instability-v0.1
clinical-fluid-balance-instability-v0.1
What this dataset does
This dataset evaluates whether models can detect instability in fluid balance dynamics.
Each row represents a simplified clinical fluid-management scenario observed across three time points.
The task is to determine whether the system remains volume-stable or is moving toward fluid overload instability.
Core stability idea
Fluid instability does not depend on fluid input alone.
A patient may receive… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-fluid-balance-instability-v0.1.
