datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
met-office-uk-deterministic-solar
Met Office UK Deterministic Dataset (Zarr Format)
Description
This dataset is a subset of the Met Office UK Deterministic Dataset, converted from the original NetCDF format into Zarr format for modern data analysis. The Zarr files are packaged as .zarr.zip archives for efficient storage and transfer.
The subset focuses on specific variables and configurations, which are detailed below. Researchers and developers can use this subset for applications in climate science… See the full description on the dataset page: https://huggingface.co/datasets/openclimatefix/met-office-uk-deterministic-solar.SolarWM-Data
SolarWM-Data
SolarWM-Data is a reusable video-data foundation for camera-conditioned
world-model research. The main Hugging Face repository publishes portable
release controls, licenses, deterministic test indexes, and directly readable
format examples. It also contains the SolarWM-Data-Annotation/
reconstruction package. The full raw video and preencoded latent payloads are
distributed separately because of their size and upstream terms.
Project Page: SolarWM
SolarWM-Data/… See the full description on the dataset page: https://huggingface.co/datasets/junchaoh-cs/SolarWM-Data.solarchive
solarchive.org: Solana Blockchain Datasets
A clean, long-term, public archive of Solana blockchain data.
This dataset contains a complete historical archive of Solana blockchain transactions, accounts, and tokens, sourced from Google BigQuery's public Solana dataset and optimized for analysis.
🎯 What is this?
Solarchive is a free, public archive of the entire Solana blockchain, designed for:
🔬 Researchers analyzing blockchain behavior and patterns
📊 Data scientists… See the full description on the dataset page: https://huggingface.co/datasets/solarchive/solarchive.SKIPPD
Citation
If you find SKIPP'D useful to your research, please cite:
Nie, Y., Li, X., Scott, A., Sun, Y., Venugopal, V., & Brandt, A. (2023). SKIPP’D: A SKy Images and Photovoltaic Power Generation Dataset for short-term solar forecasting. Solar Energy, 255, 171-179.
or
@article{nie2023skipp,
title={SKIPP’D: A SKy Images and Photovoltaic Power Generation Dataset for short-term solar forecasting},
author={Nie, Yuhao and Li, Xiatong and Scott, Andea and Sun, Yuchi and Venugopal… See the full description on the dataset page: https://huggingface.co/datasets/solarbench/SKIPPD.solar-flare-hmi-datasetSolaris_parquetsolar-system-moons
Solar System Moons
Credit: NASA/JPL-Caltech
Part of a dataset collection on Hugging Face.
Dataset description
Every known natural satellite of planets and dwarf planets in the Solar System with orbital elements, physical parameters, and discovery data. Sourced from NASA JPL Solar System Dynamics.
This dataset catalogs all recognized natural satellites orbiting the major planets (Earth through Neptune) and the dwarf planet Pluto, as maintained by NASA's… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/solar-system-moons.SolarFlowRefiner-Datasetsolaris-eval-datasets
Solaris Eval Datasets
Project Page | Paper | Github
Evaluation datasets collected via SolarisEngine to evaluate the Solaris multiplayer world model for Minecraft. Refer to Solaris repository for downloading and evaluation running instructions.
You can also use it to benchmark any multi-agent action-conditioned video model.
Dataset Info
The dataset contains videos and actions for two players in Minecraft at 720p and 20 fps.
You can find the action space in the training… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/solaris-eval-datasets.SolarWM-Data_test-set-v1
SolarWM Standalone Test Set v1
This repository contains the complete, self-contained SolarWM test set without
the training shards. It includes 1,300 clips from 13 source views (100 per
view), packaged as 60 uncompressed WebDataset tar files totaling approximately
77.4 GB.
The test identities are the current accepted SolarWM standalone evaluation
split, excluding MIND. The release contains 757 xhigh and 543 high samples. All selected
samples have non-empty captions and finite… See the full description on the dataset page: https://huggingface.co/datasets/junchaoh-cs/SolarWM-Data_test-set-v1.solar-pv-detection-brandenburg-dataset
Solar PV Ground Truth Dataset — Brandenburg Orthophotos
Hand-corrected ground-truth masks for photovoltaic detection on 20 cm GSD
aerial orthophotos (1 km × 1 km tiles, 4-band RGBI) of Brandenburg, Germany.
Companion to the model repository
solar-pv-segmentation-brandenburg.
Contents
Folder
Contents
gt_masks_selected/
1432 hand-corrected patch masks (256×256 px, uint8, 0=background / 1=PV), split into training/validation/testing
gt_masks_full_tiles/… See the full description on the dataset page: https://huggingface.co/datasets/Muemmel/solar-pv-detection-brandenburg-dataset.curvature-repro-resultssolaris-training-dataset
Solaris Training Dataset
Project Page | Paper | Github
The training dataset collected via SolarisEngine to train the Solaris multiplayer world model for Minecraft.
Dataset Info
The dataset contains 12.64 M frames across two players. It was recorded at 20 FPS. The observation (video) dimensions are: 1280 × 720.
The action space is presented below:
Action key
Type
Description
forward
bool/sustainedPlayer moving forward (W).
back
bool/sustained
Player… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/solaris-training-dataset.details_fblgit__UNA-SOLAR-10.7B-Instruct-v1.0
Dataset Card for Evaluation run of fblgit/UNA-SOLAR-10.7B-Instruct-v1.0
Dataset automatically created during the evaluation run of model fblgit/UNA-SOLAR-10.7B-Instruct-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_fblgit__UNA-SOLAR-10.7B-Instruct-v1.0.AgentBankf107-solar-flux
F10.7 Solar Radio Flux (Penticton)
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Daily F10.7 cm (2800 MHz) solar radio flux measurements from the Dominion Radio Astrophysical Observatory in Penticton, BC. The primary proxy for solar extreme ultraviolet (EUV) radiation, measured continuously since 1947.
The F10.7 solar radio flux is THE primary proxy for solar extreme ultraviolet (EUV) radiation. It has been measured… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/f107-solar-flux.solar-wind
Real-Time Solar Wind (DSCOVR/ACE)
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Real-time solar wind plasma and magnetic field measurements from the DSCOVR and ACE spacecraft at the L1 Lagrange point, via NOAA SWPC. Updated daily.
The solar wind is a continuous stream of charged particles flowing from the Sun. Its speed, density, and magnetic field orientation (especially Bz) are the primary drivers of geomagnetic storms.… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/solar-wind.Surya-bench-solarwind
Solar Wind Forecasting Dataset
Dataset Summary
This dataset provides hourly solar wind plasma and interplanetary magnetic field (IMF) parameters at L1, derived from NASA’s OMNI dataset. The primary forecasting target is the solar wind speed (V), while additional parameters are included for completeness:
Solar wind speed (V)
IMF Bx (GSE)
IMF By (GSM)
IMF Bz (GSM)
Proton number density (N)
The dataset is structured for machine learning experiments, particularly… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Surya-bench-solarwind.solar-flare-events
Solar Flare Events (GOES X-ray)
Credit: NASA/SDO
Part of a dataset collection on Hugging Face.
Dataset description
Individual solar flare detections from GOES X-ray sensors (2017-present) with class, peak flux, and timing. Updated daily.
Solar flares are sudden bursts of electromagnetic radiation from the Sun. They are classified by peak X-ray flux in the 1-8 Angstrom band: B (< 10^-6 W/m2), C (10^-6), M (10^-5), and X (10^-4 W/m2). M and X-class flares… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/solar-flare-events.solar-MID-descriptorsFuseChat-Mixture-NH-2-SOLAR-10.7B-Representation
Dataset Card for FuseChat-Mixture
Dataset Description
FuseChat-Mixture is the training dataset used in 📑FuseChat: Knowledge Fusion of Chat Models
FuseChat-Mixture is a comprehensive training dataset covers different styles and capabilities, featuring both human-written and model-generated, and spanning general instruction-following and specific skills. These sources include:
Orca-Best: We sampled 20,000 examples from Orca-Best, which is filtered from the… See the full description on the dataset page: https://huggingface.co/datasets/FuseAI/FuseChat-Mixture-NH-2-SOLAR-10.7B-Representation.solar-radio-bursts
Solar Radio Burst Events
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Catalog of solar radio burst events (spectral sweeps, fixed-frequency bursts, noise storms) from NOAA SWPC. Updated daily with incremental merge.
Solar radio bursts are produced by energetic electrons accelerated during solar flares and coronal mass ejections. They are important indicators of space weather activity:
Spectral sweeps (RSP) —… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/solar-radio-bursts.details_LDCC__LDCC-SOLAR-10.7B
Dataset Card for Evaluation run of LDCC/LDCC-SOLAR-10.7B
Dataset automatically created during the evaluation run of model LDCC/LDCC-SOLAR-10.7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_LDCC__LDCC-SOLAR-10.7B.solar-panel-inspectiondetails_dddsaty__SOLAR-Instruct-ko-Adapter-Attach
Dataset Card for Evaluation run of dddsaty/SOLAR-Instruct-ko-Adapter-Attach
Dataset automatically created during the evaluation run of model dddsaty/SOLAR-Instruct-ko-Adapter-Attach on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_dddsaty__SOLAR-Instruct-ko-Adapter-Attach.details_upstage__SOLAR-10.7B-v1.0
Dataset Card for Evaluation run of upstage/SOLAR-10.7B-v1.0
Dataset automatically created during the evaluation run of model upstage/SOLAR-10.7B-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_upstage__SOLAR-10.7B-v1.0.hk_content_corpus
HK Content Corpus (Cantonese & Traditional Chinese)
This dataset contains eight cleaned source-specific corpora of Hong Kong Cantonese and Traditional Chinese text, crawled from public websites and platforms.
It was initially created for the experiments reported in https://doi.org/10.1145/3744341 which study the effect of diglossia on Hong Kong language modeling.
Each file stores plain UTF-8 text, where each record occupies one line, and blank lines serve as separators.
This… See the full description on the dataset page: https://huggingface.co/datasets/SolarisCipher/hk_content_corpus.3_4_fusechat_v1_openchat-3.5_mixtral-8x7b-instruct-v0.1_solar-10.7b-instruct-v1.0_representationdetails_jilp00__Hermes-2-SOLAR-10.7B-Symbolic
Dataset Card for Evaluation run of jilp00/Hermes-2-SOLAR-10.7B-Symbolic
Dataset automatically created during the evaluation run of model jilp00/Hermes-2-SOLAR-10.7B-Symbolic on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_jilp00__Hermes-2-SOLAR-10.7B-Symbolic.SOLAR
SOLAR — ARCLE trajectories for ARC, written by an LLM
Synthesized Offline Learning dataset for Abstraction and Reasoning
Step-by-step solutions to the 400 tasks of the ARC-AGI-1 training split, on
inputs resampled by RE-ARC rather than
the original pairs, executable in ARCLE.
Every trajectory is a sequence of grid operations — select a region, recolor it,
move an object, paste, submit — recorded with the full environment state at each
step, so it can be replayed, watched… See the full description on the dataset page: https://huggingface.co/datasets/dbsgh797210/SOLAR.
