datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fusion-dwfire-fusion-wa-1000m
FireFusion WA 1000m
Daily spatio-temporal datacube for wildfire ignition and cause prediction over Washington State, an envelope spanning the Puget lowlands east to the Idaho border. Ten geospatial products spanning terrain, fuels, weather, human activity, lightning, and fire history are aggregated onto a single daily 1km by 1km grid covering every fire season 2003-2020.
Daily fire-season coverage, May 1 - Oct 31 of every year 2003-2020; the window contains every recorded… See the full description on the dataset page: https://huggingface.co/datasets/torq1/fire-fusion-wa-1000m.constellaration
Dataset Card for ConStellaration
A dataset of diverse quasi-isodynamic (QI) stellarator boundary shapes with corresponding performance metrics and ideal magneto-hydrodynamic (MHD) equilibria, as well as settings for their generation.
The performance metrics and ideal MHD equilibria were evaluated under vacuum (default) and with plasma inside (finite beta).
Dataset Details
Dataset Description
Stellarators are magnetic confinement devices that are… See the full description on the dataset page: https://huggingface.co/datasets/proxima-fusion/constellaration.fusion-equilibrium-challenge
⚛️ The Fusion Equilibrium Challenge
Predict the shape of a fusion plasma (the magnetic equilibrium, $\psi$) from
control inputs and diagnostics alone — without magnetic sensors. Data comes from
two tokamaks:
DIII-D — General Atomics tokamak (San Diego, USA)
MAST — Mega Ampere Spherical Tokamak (Culham, UK)
In data-science terms this is an image-regression / control problem: predict a
2-D poloidal flux map (efit_psirz) at each EFIT timestep from coil currents
(the actuators)… See the full description on the dataset page: https://huggingface.co/datasets/Sophelio/fusion-equilibrium-challenge.fire-fusion-wa-4000m
FireFusion WA 4000m
Daily spatio-temporal datacube for wildfire ignition and cause prediction over Washington State, an envelope spanning the Puget lowlands east to the Idaho border. Ten geospatial products spanning terrain, fuels, weather, human activity, lightning, and fire history are aggregated onto a single daily 4km by 4km grid covering every fire season 2003-2020.
Daily fire-season coverage, May 1 - Oct 31 of every year 2003-2020; the window contains every recorded… See the full description on the dataset page: https://huggingface.co/datasets/torq1/fire-fusion-wa-4000m.autonomous-vehicle-sensor-fusion
Autonomous Vehicle Raw Sensor Telemetry & Simulation Dataset
This repository contains raw, uncompressed sensor buffers, massive neural network weights, federated learning node dumps, and VRAM memory snapshots collected from autonomous vehicle test fleets. Data is provided "as is" for offline perception model training, system debugging, and Hardware-in-the-Loop (HIL) simulations.
Dataset Structure (Full Schema)
Due to horizontal scaling and massive daily ingestions… See the full description on the dataset page: https://huggingface.co/datasets/born5149/autonomous-vehicle-sensor-fusion.constellaration-bench-submissionsfire-fusion-wa-2000m
FireFusion WA 2000m
Daily spatio-temporal datacube for wildfire ignition and cause prediction over Washington State, an envelope spanning the Puget lowlands east to the Idaho border. Ten geospatial products spanning terrain, fuels, weather, human activity, lightning, and fire history are aggregated onto a single daily 2km by 2km grid covering every fire season 2003-2020.
Daily fire-season coverage, May 1 - Oct 31 of every year 2003-2020; the window contains every recorded… See the full description on the dataset page: https://huggingface.co/datasets/torq1/fire-fusion-wa-2000m.constellaration-bench-resultsFUSION-Finetune-12M
FUSION-12M Dataset
Please see paper & website for more information:
https://arxiv.org/abs/2504.09925
https://github.com/starriver030515/FUSION
Overview
FUSION-12M is a large-scale, diverse multimodal instruction-tuning dataset used to train FUSION-3B and FUSION-8B models. It builds upon Cambrian-1 by significantly expanding both the quantity and variety of data, particularly in areas such as OCR, mathematical reasoning, and synthetic high-quality Q&A data. The goal is… See the full description on the dataset page: https://huggingface.co/datasets/starriver030515/FUSION-Finetune-12M.fusionX_480p_wan21_latentsFUSION-Pretrain-10M
FUSION-10M Dataset
Please see paper & website for more information:
https://arxiv.org/abs/2504.09925
https://github.com/starriver030515/FUSION
Overview
FUSION-10M is a large-scale, high-quality dataset of image-caption pairs used to pretrain FUSION-3B and FUSION-8B models. It builds upon established datasets such as LLaVA, ShareGPT4, and PixelProse. In addition, we synthesize 2 million task-specific image-caption pairs to further enrich the dataset. The goal of… See the full description on the dataset page: https://huggingface.co/datasets/starriver030515/FUSION-Pretrain-10M.cmu_play_fusion_rawfire-fusion-cascades-500m
FireFusion Cascades 500m
Daily spatio-temporal datacube for wildfire ignition and cause prediction over the Eastern Cascades of Washington State, a 272 km square running from the Cascade crest through the Okanogan Highlands, the most fire-active terrain in the state. Ten geospatial products spanning terrain, fuels, weather, human activity, lightning, and fire history are aggregated onto a single daily 500m by 500m grid covering every fire season 2003-2020.
Daily fire-season… See the full description on the dataset page: https://huggingface.co/datasets/torq1/fire-fusion-cascades-500m.cmu_play_fusion_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "franka",
"total_episodes": 576,
"total_frames": 235922,
"total_tasks": 44,
"total_videos": 576,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:576"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/cmu_play_fusion_lerobot.adapter-based-multimodal-fusion
Falcon-Audio Training Dataset
Training-ready Parquet shards for Falcon-Audio. Rows contain Gemma-tokenized inputs/labels and fp16 Whisper encoder features encoded as raw bytes.
NoReGeo
NoReGeo: Non-Reasoning Geometry Benchmark
Evaluation samples from NoReGeo benchmark. Each problem is shown in three formats -- (a) text‑only, (b) text with dotted-image (points only), and (c) text with full-image (points plus connecting lines) -- together with the golden answer (yellow) and the model’s prediction.
NoReGeo is a benchmark for evaluating intrinsic geometric understanding in LLMs—without reasoning, algebra, or chain-of-thought. It contains 2,500 trivial… See the full description on the dataset page: https://huggingface.co/datasets/FusionBrainLab/NoReGeo.FusionX-Multimodal-Sample-Data-V3
FusionX Multimodal Sample Dataset (V3)
This repository contains a multimodal dataset capturing synchronized stereo vision, RGB, IMU, and tactile glove data across 11 distinct tasks. It is intended for research in multimodal perception, manipulation learning, and tactile-vision fusion.
Dataset Overview
Each task is captured with the following synchronized modalities:
Mono Stereo Vision — Left and right monochrome camera streams stored as raw .png files at 640×400… See the full description on the dataset page: https://huggingface.co/datasets/touchtronix/FusionX-Multimodal-Sample-Data-V3.asd
Musubi Tuner
English | 日本語
Click to expand
Musubi Tuner
Table of Contents
Introduction
Sponsors
Support the Project
Recent Updates
Releases
For Developers Using AI Coding Agents
Overview
Hardware RequirementsFeatures
Documentation
Installation
pip based installation
uv based installation
Linux/MacOS
Windows
Model Download
Usage
Dataset Configuration
Pre-caching and Training
Configuration of Accelerate
Training and Inference
Miscellaneous
SageAttention Installation
PyTorch… See the full description on the dataset page: https://huggingface.co/datasets/FusionCow/asd.fusionX_480pfusion-afdb-quality-real
FusionUncertaintyNet — Real AFDB Quality Dataset
298,547 real proteins harvested from AlphaFold DB (v6 models).
sequence: UniProt reviewed (Swiss-Prot), length 30–1022 aa
plddt[]: REAL per-residue pLDDT parsed from AFDB PDB B-factors (CA atoms)
target[]: lDDT-style quality = pLDDT − disorder·10 (documented proxy), clipped 1–100
phi[]/psi[]: Ramachandran-basin priors (α/β gaussians)
Sharded as manifest_shard_NNN.jsonl (~4.8k rows each)
Provenance: rest.uniprot.org stream +… See the full description on the dataset page: https://huggingface.co/datasets/bhumika-tewari-282006/fusion-afdb-quality-real.FusionAudiofusion360_test_meshFusionSensecmu_play_fusionThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 576,
"total_frames": 235922,
"total_tasks": 44,
"total_videos": 576,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:576"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/cmu_play_fusion.survival-fusion-tcga600-v3.3
Survival Fusion TCGA-600 v3.3
This dataset contains reproducibly derived multimodal features for 600 TCGA
patients: 200 each from TCGA-BLCA, TCGA-HNSC, and TCGA-STAD. It supports the
Survival Fusion benchmark's development and frozen Stage-1 qualification
protocol.
The repository contains derived feature arrays and public TCGA identifiers. It
does not contain raw whole-slide images or raw sequencing reads.
Contents
canonical-v3.3-v1/: canonical per-patient… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/survival-fusion-tcga600-v3.3.lm-eval-results-yunconglong-Truthful_DPO_TomGrc_FusionNet_7Bx2_MoE_13B-private
Dataset Card for Evaluation run of yunconglong/Truthful_DPO_TomGrc_FusionNet_7Bx2_MoE_13B
Dataset automatically created during the evaluation run of model yunconglong/Truthful_DPO_TomGrc_FusionNet_7Bx2_MoE_13B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-yunconglong-Truthful_DPO_TomGrc_FusionNet_7Bx2_MoE_13B-private.drifting-vla-v2-cmu_play_fusionfusion-pairwise-evals-test-time-scaling
Automatic pairwise preference evaluations for: Making, not taking, the Best-of-N
Content
This data contains pairwise automatic win-rate evaluations for the m-ArenaHard-v2.0 benchmark and it compares CommandA against gemini-2.5-pro in 2 settings:
Test-time scaling with Fusion : 5 samples are generated from CommandA, then fused with CommandA into one completion and compared to a single completion from gemini-2.5-pro
Test-time scaling with BoN : 5 samples are generated from… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/fusion-pairwise-evals-test-time-scaling.fusion360_test_scan
