datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
windsorml
WindsorML: High-Fidelity Computational Fluid Dynamics dataset for automotive aerodynamics
Contact:
Neil Ashton (NVIDIA) - contact@caemldatasets.org
website:
https://caemldatasets.org
Summary:
This work presents a new open-source high-fidelity dataset for Machine Learning (ML) containing 355 geometric variants of the Windsor body, to help the development and testing of ML surrogate models for external automotive aerodynamics. Each… See the full description on the dataset page: https://huggingface.co/datasets/neashton/windsorml.fractal_rawwindtunnel-20k
Wind Tunnel Dataset
The Wind Tunnel Dataset contains 19,812 OpenFOAM simulations of 1,000 unique automobile-like objects placed in a virtual wind tunnel measuring 20 meters long, 10 meters wide, and 8 meters high.
Each object was tested under 20 different conditions: 4 random wind speeds ranging from 10 to 50 m/s, and 5 rotation angles (0°, 180° and 3 random angles).
The object meshes were generated using Instant Mesh based on images sourced from the Stanford Cars Dataset. To… See the full description on the dataset page: https://huggingface.co/datasets/inductiva/windtunnel-20k.FineVision
Fine Vision
FineVision is a massive collection of datasets with 17.3M images, 24.3M samples, 88.9M turns, and 9.5B answer tokens, designed for training state-of-the-art open Vision-Language-Models.
More detail can be found in the blog post: https://huggingface.co/spaces/HuggingFaceM4/FineVision
Load the data
from datasets import load_dataset, get_dataset_config_names
# Get all subset names and load the first one
available_subsets =… See the full description on the dataset page: https://huggingface.co/datasets/Windwave/FineVision.windows_osworld_file_cache
OSWorld File Cache
This repository serves as a file cache for the OSWorld project, providing reliable and fast access to evaluation files that were previously hosted on Google Drive.
Overview
OSWorld is a scalable, real computer environment for multimodal agents, supporting task setup, execution-based evaluation, and interactive learning across various operating systems and applications. This cache repository ensures that all evaluation files are consistently accessible… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/windows_osworld_file_cache.Handwritten-Latex-Datasets
Dataset
This data set includes common handwritten formulas in junior high schools and high schools, and is labeled in Latex format. Can be used to train models that recognize common numbers, fractions, and sets.
Dataset source
Collected in various junior high schools and high schools, handwritten by students.
Usage
The label is stored at json folder and scanned hand-writted pictures are stored at pic folder.
Scan the qr code of the picture to get the index and… See the full description on the dataset page: https://huggingface.co/datasets/WindyVerse/Handwritten-Latex-Datasets.pl-kwsWinDeskGroundA Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments
WinDeskGround is a reproducible benchmark toolkit for evaluating GUI grounding models in realistic multi-window desktop environments.
It includes:
sensitivity-controlled experiments for occlusion, semantic distraction, and clutter
simulation difficulty levels from L1 to L5
unified evaluation scripts for multiple local or API-based models
MultiAgent-Content
MultiAgent UE5 Content
MultiAgent-Unreal 项目的配套 UE5 资产库。
快速克隆
#### 安装 Git LFS ####
## Mac 系统
brew install git-lfs
git lfs install
## Linux 系统
sudo apt-get install git-lfs
git lfs install
#### 克隆到项目目录 ####
cd unreal_project
## 有🪜
git clone https://huggingface.co/datasets/WindyLab/MultiAgent-Content Content
## 中国用户
git clone https://hf-mirror.com/datasets/WindyLab/MultiAgent-Content Content
📂 目录结构
Content/
├── Agent/ # 智能体配置
├──… See the full description on the dataset page: https://huggingface.co/datasets/WindyLab/MultiAgent-Content.WindowsOfficial repository link: microsoft/SWE-bench-Live
Paper link: https://arxiv.org/abs/2603.05026
Developing projects compatible on Windows platform is important to expand the user market.
There are some bugs that would only occur on Windows. To migrate projects to Windows some codes need to be rewritten to have branching points across different os or use cross-platform compatible libraries...
To test LLM's knowledge of Windows-specific SWE knowledge AND powershell terminal operation capability… See the full description on the dataset page: https://huggingface.co/datasets/SWE-bench-Live/Windows.eis-subset50
EIS-Subset50: 1970s U.S. Environmental Impact Statements
A 50-document test subset of scanned 1970s U.S. federal Environmental Impact
Statements (EIS) from the Northwestern University Library collection, built to
test whether current models can handle this kind of data: very long documents
(53–590 pages each), degraded microfilm scans, dense bureaucratic text, and
figure-heavy sections (maps, route corridors, site plans, photos).
Contents at a glance
Component
Size
What it… See the full description on the dataset page: https://huggingface.co/datasets/Windsao/eis-subset50.OpenGPT-4o-Image
OpenGPT-4o-Image Dataset
We introduce OpenGPT-4o-Image, a large-scale dataset constructed using a novel methodology that combines hierarchical task taxonomy with automated data generation. Our taxonomy not only includes fundamental capabilities such as text rendering and style control but also introduces highly practical yet challenging categories like scientific imagery for chemistry illustrations and complex instruction editing requiring simultaneous execution of multiple… See the full description on the dataset page: https://huggingface.co/datasets/WINDop/OpenGPT-4o-Image.wind_farms_minutely
wind_farms_minutely (TsFile format)
Minutely time series representing the wind power production of 339 wind farms in Australia.
This repository contains the full source .tsf series from the Monash Time Series Forecasting Repository converted to Apache TsFile format.
Summary
Source dataset: Monash-University/monash_tsf
Original source: https://zenodo.org/record/4654909
Monash subset: wind_farms_minutely
Modalities: Time-series
Source series: 339
Rows: 172,178,060… See the full description on the dataset page: https://huggingface.co/datasets/THULab/wind_farms_minutely.windowswindows_osworld
Dataset Card for Dataset Name
This repository contains the task examples, retrieval documents (in the archive evaluation_examples.zip), and virtual machine snapshots for benchmark OSWorld (loaded by VMware/VirtualBox depending on the machine architecture x86 or arm64).
You can find more information from our paper OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
paper Arxiv link: https://arxiv.org/abs/2404.07972
project website:… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/windows_osworld.libero_90_allThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 3917,
"total_frames": 567494,
"total_tasks": 73,
"total_videos": 0,
"total_chunks": 4,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:3917"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/libero_90_all.solar-wind
Real-Time Solar Wind (DSCOVR/ACE)
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Real-time solar wind plasma and magnetic field measurements from the DSCOVR and ACE spacecraft at the L1 Lagrange point, via NOAA SWPC. Updated daily.
The solar wind is a continuous stream of charged particles flowing from the Sun. Its speed, density, and magnetic field orientation (especially Bz) are the primary drivers of geomagnetic storms.… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/solar-wind.libero_90_all_and_metadataThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 3917,
"total_frames": 567494,
"total_tasks": 73,
"total_videos": 0,
"total_chunks": 4,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:3917"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/libero_90_all_and_metadata.dsrl-data-rotatedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 3651,
"total_frames": 399524,
"total_tasks": 67,
"total_videos": 0,
"total_chunks": 4,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:3651"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/dsrl-data-rotated.base-data-rotatedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 2443,
"total_frames": 385939,
"total_tasks": 57,
"total_videos": 0,
"total_chunks": 3,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:2443"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/base-data-rotated.extreme_randomization_6_brick_03This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 2108,
"total_frames": 419545,
"total_tasks": 1,
"total_videos": 4216,
"total_chunks": 3,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:2108"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/extreme_randomization_6_brick_03.MfS35KZeroStereo: Zero-shot Stereo Matching from Single Images
lib90-human-rotate2-part1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 61,
"total_frames": 14412,
"total_tasks": 61,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:61"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/lib90-human-rotate2-part1.scripted_atomic_train_frac_0.3_large_goal_annotationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 664,
"total_frames": 116214,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:664"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_train_frac_0.3_large_goal_annotation.china-a-share-zipts-satfire-processed-windowpudding_uncondThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 300,
"total_frames": 107260,
"total_tasks": 2,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:300"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/pudding_uncond.WindowsServer_21H2scripted_atomic_step_pose_0.6This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 955,
"total_frames": 159935,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:955"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_step_pose_0.6.scripted_atomic_step_train_frac0.3_large_imageThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 664,
"total_frames": 116214,
"total_tasks": 1,
"total_videos": 1328,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:664"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_step_train_frac0.3_large_image.
