datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FluidThermal-CUA
flowstate-cua
computer-use data for fluid and thermal simulation analysis.
what this dataset represents
flowstate-cua connects scientific tasks to real desktop screenshots, executed actions, and results that can be checked against the underlying physical fields. it provides 8,260 demonstrations and 269,058 recorded interactions for training and evaluating agents that use scientific software.
the physical fields and task instructions are synthetic. the screenshots… See the full description on the dataset page: https://huggingface.co/datasets/harrrshall/FluidThermal-CUA.OpenSeeSimE-Fluid
OpenSeeSimE-Fluid: Engineering Simulation Visual Question Answering Benchmark
Dataset Summary
OpenSeeSimE-Fluid is a large-scale benchmark dataset for evaluating vision-language models on computational fluid dynamics (CFD) simulation interpretation tasks. It contains approximately 98,000 question-answer pairs across parametrically-varied fluid simulations including turbulent flow, heat transfer, and complex flow patterns.
Purpose
While vision-language models… See the full description on the dataset page: https://huggingface.co/datasets/cmudrc/OpenSeeSimE-Fluid.OpenSeeSimE-Fluid-Small
OpenSeeSimE-Fluid-Small
A stratified 10% subset of cmudrc/OpenSeeSimE-Fluid for evaluating vision-language models at a reduced compute footprint while preserving the joint distribution of simulation type, question type, media type, and question id.
Subset Provenance
Parent dataset: cmudrc/OpenSeeSimE-Fluid (98,326 rows total)
Rows in this subset: 9,881 (10.05% of parent)
Source classes: Bent Pipe, Converging Nozzle, Heat Exchanger, Heat Sink, Mixing Pipe
Parquet shards:… See the full description on the dataset page: https://huggingface.co/datasets/cmudrc/OpenSeeSimE-Fluid-Small.OpenSeeSimE-Fluid-Mini
OpenSeeSimE-Fluid-Mini
A stratified 1% subset of cmudrc/OpenSeeSimE-Fluid for evaluating vision-language models at a reduced compute footprint while preserving the joint distribution of simulation type, question type, media type, and question id.
Subset Provenance
Parent dataset: cmudrc/OpenSeeSimE-Fluid (98,326 rows total)
Rows in this subset: 1,040 (1.06% of parent)
Source classes: Bent Pipe, Converging Nozzle, Heat Exchanger, Heat Sink, Mixing Pipe
Parquet shards: 3… See the full description on the dataset page: https://huggingface.co/datasets/cmudrc/OpenSeeSimE-Fluid-Mini.simple-fluid-simulationsfluid-2-sft-eval
Fluid 2 — dictation cleanup eval
This repository contains 7,161 text-only voice-dictation cleanup evaluation
rows for like-for-like model comparison. No audio is included or fetched.
Split
Rows
Documents
Audio
eval
7,161
2,718
Not included
What this benchmark tests
The benchmark measures whether a model can turn noisy voice dictation into the
intended written text without answering it or adding content. The rows cover:
local ASR, spelling… See the full description on the dataset page: https://huggingface.co/datasets/johnbean393/fluid-2-sft-eval.OpenSeeSimE-Fluid-Testfluid-2-sft-asr
Fluid 2 — synthetic dictation cleanup
Fluid 2 is an English supervised-fine-tuning corpus for models that turn noisy automatic-speech-recognition output into the written insertion a user intended. It contains 354,549 rows in official document-grouped 96/2/2 splits, 861.3 hours of processed 16 kHz speech, and 8.48M target-side loss tokens in 355 Parquet shards (49.25 GiB).
This is not an ordinary transcription dataset. The model sees document context plus an ASR hypothesis and… See the full description on the dataset page: https://huggingface.co/datasets/johnbean393/fluid-2-sft-asr.
