datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fluxloracurated-danbooru-2026-512px-flux2-vaemjnj_flux32curated-danbooru-2026-256px-flux2-vaef107-solar-flux
F10.7 Solar Radio Flux (Penticton)
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Daily F10.7 cm (2800 MHz) solar radio flux measurements from the Dominion Radio Astrophysical Observatory in Penticton, BC. The primary proxy for solar extreme ultraviolet (EUV) radiation, measured continuously since 1947.
The F10.7 solar radio flux is THE primary proxy for solar extreme ultraviolet (EUV) radiation. It has been measured… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/f107-solar-flux.goes-xray-flux
GOES Solar X-Ray Flux (1-Minute)
Credit: NASA/SDO
Part of a dataset collection on Hugging Face.
Dataset description
Solar soft X-ray flux from the GOES X-Ray Sensor (XRS), the operational backbone of solar flare monitoring. Updated daily from NOAA SWPC, growing incrementally at 1-minute cadence.
The GOES (Geostationary Operational Environmental Satellite) X-Ray Sensor measures the Sun's soft X-ray irradiance in two wavelength bands: a "short" 0.05-0.4 nm… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/goes-xray-flux.iras-sky-flux-plates
IRAS Sky Flux Plates
IRAS high-resolution hours-confirmed survey intensity and statistical-weight plates in the four survey bands.
Data structure
Each represented .INTE or .STAT archive member is one source-named configuration. Raw FITS axes and any whole-pixel truncation are retained.
The holding has 4,827 source-named configurations from 212 served archives containing 4,832 members.
Raw values, source order, FITS image axes and structured provenance are… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/iras-sky-flux-plates.goes_omni_electron_flux_forecasting
GOES–OMNI >2 MeV Electron Flux Forecasting
This dataset combines cross-calibrated NOAA GOES-14/GOES-16 >2 MeV electron
flux with NASA/GSFC OMNI solar-wind and geomagnetic drivers on a uniform
five-minute UTC grid.
It provides two configurations:
ml-ready (default): scaled causal features, validity flags, unscaled
30-minute/6-hour/12-hour targets, and leakage-safe chronological splits.
scientific-master: unscaled source measurements, instrument context,
calibration factors, and… See the full description on the dataset page: https://huggingface.co/datasets/THULab/goes_omni_electron_flux_forecasting.Time-Series-Library
Time-Series-Library (TSLib)
TSLib is an open-source library for deep learning researchers, especially for deep time series analysis.
We provide a neat code base to evaluate advanced deep time series models or develop your model, which covers five mainstream tasks: long- and short-term forecasting, imputation, anomaly detection, and classification.
This benchmark collection is designed to evaluate and develop advanced deep time-series models. For an in-depth exploration of… See the full description on the dataset page: https://huggingface.co/datasets/fluxae/Time-Series-Library.mjnj_640_flux2top-flutter-packages
Top Flutter Packages Dataset
Flutter is an open source framework by Google for building beautiful, natively compiled, multi-platform applications from a single codebase. It is gaining quite a bit of popularity because of ability to code in a single language and have it running on Android/iOS and web as well.
This dataset contains a snapshot of Top 5000+ flutter/dart packages hosted on Flutter package repository
The dataset was scraped in August-2024.
We aim to use this dataset to… See the full description on the dataset page: https://huggingface.co/datasets/deepklarity/top-flutter-packages.ds1234_flux32flutter-diff-steps-v1
Flutter Codegen: Diff Steps
Synthetic dataset of step-by-step Flutter/Dart widget construction, where each
row is one incremental edit in a sequence: given a goal, the current code, and the
history of steps taken so far, predict the next action (a short description) and
the code change as a search/replace diff hunk.
Built for training and evaluating small language models on iterative, diff-based
code editing -- as opposed to regenerating the whole file at each step. This is
the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/flutter-diff-steps-v1.flux2-klein-latent-trajectories
FLUX.2 Klein Latent Trajectories
This dataset contains synthetic text-to-image generations produced with a local FLUX.2 Klein 4B snapshot, together with the prompts, generated WebP images, and captured intermediate diffusion latent states.
Contents
262,144 generated examples.
512 Torch shard files under shards/, with 512 samples per shard.
One JSON sidecar per shard with generation metadata.
Image resolution: 512 x 512.
Image format: WebP, quality 90.
Diffusion… See the full description on the dataset page: https://huggingface.co/datasets/ryanhlewis/flux2-klein-latent-trajectories.cv-corpus-25.0-ja
Mozilla Common Voice 25.0 - Japanese Test Set (Complete)
Dataset Description
Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace.
Key Features
Size: 9,019 validated test utterances
Coverage: 100% of official Common Voice 25.0 Japanese test split
Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.fluorescencegoes-omni-electron-flux-forecasting
GOES–OMNI >2 MeV Electron Flux Forecasting
This dataset combines cross-calibrated NOAA GOES-14/GOES-16 >2 MeV electron
flux with NASA/GSFC OMNI solar-wind and geomagnetic drivers on a uniform
five-minute UTC grid.
It provides two configurations:
ml-ready (default): scaled causal features, validity flags, unscaled
30-minute/6-hour/12-hour targets, and leakage-safe chronological splits.
scientific-master: unscaled source measurements, instrument context,
calibration factors, and… See the full description on the dataset page: https://huggingface.co/datasets/snowsadh/goes-omni-electron-flux-forecasting.ds234_flux32Rainbow-Pony-100m-Flutter-steps-eval
Rainbow-Pony-100M Flutter — Steps Mode — Validation Results
Dataset Summary
Held-out evaluation results for bbidpa/Rainbow-Pony-100m-Flutter-steps,
a 100M-parameter transformer trained from scratch to edit Flutter/Dart source files.
In steps mode, the model is given an existing file and an edit instruction and
generates a sequence of localized search/replace edit actions, each mechanically
applied to the current file state before the next action is generated… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Rainbow-Pony-100m-Flutter-steps-eval.multimodal-peer-collaboration-samples
Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles
Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges.
▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection
Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.friend-bench
Can a model — or a human — tell how two people are related from a 20-second clip of how they interact?
🌐 Built on Seamless Interaction
FriendBench is a suite of benchmarks for social perception from thin-slice dyadic
interaction — inferring facts about two people's relationship from a brief clip of how they
interact, built on the Seamless Interaction
dataset. Each released set is a config of this repository.
🎧 Multi-modal — text, audio, and video for every clip
🎯 Objective label —… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/friend-bench.flutter-full-examples-v1
Flutter Codegen: Full Examples
Synthetic dataset of complete Flutter/Dart widgets, each paired with the goal
that describes them and (optionally) starting code. Unlike flutter-codegen-diff-steps,
there's no step history or diff structure here -- each row is a single, standalone
goal -> complete file example.
This is the whole-code counterpart to flutter-diff-steps-v1, intended for
training/evaluating a baseline that generates the entire file in one shot, to
compare against the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/flutter-full-examples-v1.multimodal-expert-instruction-samples
Multimodal Expert Instruction Samples - Musical Instrument Lessons with Channel-separated Audio and Video
A music teacher and a student work through two one-on-one lessons: both voices and both instruments on separate tracks, the student on camera, with the lesson plans, the instructions given to each side and both sides' post-lesson ratings alongside.
▶ Watch the lessons · See Peer Collaboration samples · Discuss the full collection
Sister collection: Peer Collaboration… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-expert-instruction-samples.Qwen2.5-Coder-0.5B-Flutter-steps-eval
Qwen2.5-Coder-0.5B Flutter — Steps Mode — Validation Results
Dataset Summary
Held-out evaluation results for bbidpa/Qwen2.5-Coder-0.5B-Flutter-steps,
a fine-tune of Qwen2.5-Coder-0.5B for editing Flutter/Dart source files. In steps
mode, the model is given an existing file and an edit instruction and generates a
sequence of localized search/replace edit actions, each mechanically applied to the
current file state before the next action is generated, until the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Qwen2.5-Coder-0.5B-Flutter-steps-eval.tape-fluorescenceRainbow-Pony-100m-Flutter-direct-eval
Rainbow-Pony-100M Flutter — Direct Mode — Validation Results
Dataset Summary
Held-out evaluation results for bbidpa/Rainbow-Pony-100m-Flutter-direct,
a 100M-parameter transformer trained from scratch to edit Flutter/Dart source files.
In direct mode, the model is given an existing file and an edit instruction and
generates the complete modified file in a single forward pass (as opposed to the
steps / iterative diff-based mode — see the sibling dataset… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Rainbow-Pony-100m-Flutter-direct-eval.feedback-prize-english-language-learning-fluencyQwen2.5-Coder-0.5B-Flutter-direct-eval
Qwen2.5-Coder-0.5B Flutter — Direct Mode — Validation Results
Dataset Summary
Held-out evaluation results for bbidpa/Qwen2.5-Coder-0.5B-Flutter-direct,
a fine-tune of Qwen2.5-Coder-0.5B for editing Flutter/Dart source files. In direct
mode, the model is given an existing file and an edit instruction and generates the
complete modified file in a single forward pass (as opposed to the steps /
iterative diff-based mode — see the sibling dataset… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Qwen2.5-Coder-0.5B-Flutter-direct-eval.ds4_anime_flux32ConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B
Visual Memory Results: convai2-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B.
