datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
solana-dex-execution
Solana DEX quotes and routing
Jupiter swap quotes at fixed SOL input sizes, with the route legs returned for each quote. The panel supports comparisons of quoted output, reported price impact and routing across sizes and observation times.
Contents
Table
Record
solana_swap_quotes
A pair and input amount, with quoted output, price impact, value and route-leg count
solana_swap_routes
A leg of the chosen route, including venue, amounts and allocation… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/solana-dex-execution.solar-system-moons
Solar System Moons
Credit: NASA/JPL-Caltech
Part of a dataset collection on Hugging Face.
Dataset description
Every known natural satellite of planets and dwarf planets in the Solar System with orbital elements, physical parameters, and discovery data. Sourced from NASA JPL Solar System Dynamics.
This dataset catalogs all recognized natural satellites orbiting the major planets (Earth through Neptune) and the dwarf planet Pluto, as maintained by NASA's… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/solar-system-moons.solana-dex-datasolana-memecoin-calls
Solana memecoin calls — a public record with the misses left in
8,161 pump.fun token calls, each with the market cap we called it at, the peak it reached
afterwards, and the exact second it was posted publicly. The whole file is hashed and the hash is
anchored in a Bitcoin block, so no row can be added, edited or back-dated after the fact.
Every trading channel publishes its winners. This is the same feed with the losers still in it —
about six calls in ten never double, and… See the full description on the dataset page: https://huggingface.co/datasets/Smurfetc/solana-memecoin-calls.AgentBanksolar-wind
Real-Time Solar Wind (DSCOVR/ACE)
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Real-time solar wind plasma and magnetic field measurements from the DSCOVR and ACE spacecraft at the L1 Lagrange point, via NOAA SWPC. Updated daily.
The solar wind is a continuous stream of charged particles flowing from the Sun. Its speed, density, and magnetic field orientation (especially Bz) are the primary drivers of geomagnetic storms.… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/solar-wind.Surya-bench-solarwind
Solar Wind Forecasting Dataset
Dataset Summary
This dataset provides hourly solar wind plasma and interplanetary magnetic field (IMF) parameters at L1, derived from NASA’s OMNI dataset. The primary forecasting target is the solar wind speed (V), while additional parameters are included for completeness:
Solar wind speed (V)
IMF Bx (GSE)
IMF By (GSM)
IMF Bz (GSM)
Proton number density (N)
The dataset is structured for machine learning experiments, particularly… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Surya-bench-solarwind.solana-yield-honesty
Solana Honesty Index
What each Solana stablecoin product says it pays, next to what it actually
paid, measured from a share price rather than from a claim.
Snapshot generated 2026-09-25T12:21:05.434Z. Window 30 days.
13 products across 3 protocols,
13 comparable, 0 published but not
comparable. Realized figures: 5 by issuer_share_price_history, 2 by onchain_share_price, 6 by issuer_share_price_observed.
product
advertised
realized
gap
delivered
realized method
Kamino… See the full description on the dataset page: https://huggingface.co/datasets/kerne-protocol/solana-yield-honesty.Solace-1.0-Omni
Project Solace
The largest verified frontier-model distillation corpus ever released.
60 datasets · 7 frontier model families · 12,586,893 unique conversations · One file · Zero filler
The short version
This is synthetic data. The best kind of synthetic data.
Every example was generated by a verified 2026 frontier model — GLM-5.2, Claude Fable 5, Mythos 5, GPT-5.6 Sol, GPT-5.5 Codex, DeepSeek V4 Pro 0813, Qwen 3.8-Max, and Kimi K3 — then… See the full description on the dataset page: https://huggingface.co/datasets/Solstice-AI/Solace-1.0-Omni.solar-radio-bursts
Solar Radio Burst Events
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Catalog of solar radio burst events (spectral sweeps, fixed-frequency bursts, noise storms) from NOAA SWPC. Updated daily with incremental merge.
Solar radio bursts are produced by energetic electrons accelerated during solar flares and coronal mass ejections. They are important indicators of space weather activity:
Spectral sweeps (RSP) —… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/solar-radio-bursts.hk_content_corpus
HK Content Corpus (Cantonese & Traditional Chinese)
This dataset contains eight cleaned source-specific corpora of Hong Kong Cantonese and Traditional Chinese text, crawled from public websites and platforms.
It was initially created for the experiments reported in https://doi.org/10.1145/3744341 which study the effect of diglossia on Hong Kong language modeling.
Each file stores plain UTF-8 text, where each record occupies one line, and blank lines serve as separators.
This… See the full description on the dataset page: https://huggingface.co/datasets/SolarisCipher/hk_content_corpus.Solace-270K-Golden-131K-SFT
Solace-270K-Golden-131K-SFT
Official 270,000 Golden Distillation Corpus for 131K Native Context Post-Training
Executive Summary
Solstice-AI/Solace-270K-Golden-131K-SFT is the curated, high-purity post-training corpus created by Solstice-AI, extracted and balanced from the landmark 12.59M-conversation Solstice-AI/Solace-1.0-Omni foundation.
Designed specifically for 131,072 Token (131K Token) native context post-training, this dataset contains zero… See the full description on the dataset page: https://huggingface.co/datasets/Solstice-AI/Solace-270K-Golden-131K-SFT.SOLAR
SOLAR — ARCLE trajectories for ARC, written by an LLM
Synthesized Offline Learning dataset for Abstraction and Reasoning
Step-by-step solutions to the 400 tasks of the ARC-AGI-1 training split, on
inputs resampled by RE-ARC rather than
the original pairs, executable in ARCLE.
Every trajectory is a sequence of grid operations — select a region, recolor it,
move an object, paste, submit — recorded with the full environment state at each
step, so it can be replayed, watched… See the full description on the dataset page: https://huggingface.co/datasets/dbsgh797210/SOLAR.SolarChemQA_Clark
SolarChemQA
Dataset Description
SolarChemQA is a novel question answering dataset curated from solar chemistry literature designed to rigorously assess the capabilities of Large Language Models (LLMs) driven QA systems in processing domain-specific scientific content.
The dataset provides the raw extracted context from solar chemistry papers, domain expert annotations, and the domain expert validated sentences from the context may be used as evidences for the… See the full description on the dataset page: https://huggingface.co/datasets/ClarkWangPas/SolarChemQA_Clark.OpenOrca_Solar_filtered
TODO
To be consistent, we need to change column name into ['instruction', 'input', 'output'], which is same as alpaca-gpt4.
Dataset Summary
This is a filtered version of OpenOrca dataset based on Solar 10.7B paper.
In this version, of the 4.2M OpenOrca data, 113k data is removed.
In more conservative version here, of the 4.2M OpenOrca data, 117k data is removed.
Step 1
FLAN data link broken
Based on DataProvenanceInitiative/flan2021_submix_original… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/OpenOrca_Solar_filtered.wind-and-solar-candidate-olmoearth-segmentation
OlmoEarth v1.2-ready renewable energy dataset
Derived cook 20260830T191500Z from pinned source 2c551f58998cd25554ea679148b21a9c701b51db.
The branch uses no train/validation/test split. Every accepted row is tagged all.
Accepted imagery rows: 12649
Quality-unavailable rows: 4
Release mode: partial with documented quality-unavailable exclusions
Clearest-available quality fallbacks: 4
Sentinel-2 L2A: uint16 [12,1,128,128]
Band order: B02,B03,B04,B08,B05,B06,B07,B8A,B11,B12,B01,B09… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/wind-and-solar-candidate-olmoearth-segmentation.SOLAR-extras
SOLAR extras — the handcraft control, and rendered previews
Companion to dbsgh797210/SOLAR,
which holds the release itself and nothing else. Everything here is secondary to
it, which is why it is not there.
Browse the trajectories →
config
rows
tasks
what it is
handcraft
100
10
trajectories from makers written by hand, as a control
preview
400
400
one rendered picture per RE-ARC task
preview_handcraft
10
10
the same, for handcraft… See the full description on the dataset page: https://huggingface.co/datasets/dbsgh797210/SOLAR-extras.perovskite-solar-cell-efficiency-autoresearch
🔬 Perovskite Solar Cell Text Corpus for Karpathy's autoresearch
A 98.9 MB text corpus of perovskite solar cell scientific literature formatted for direct use with karpathy/autoresearch — the autonomous LLM-driven hyperparameter search framework that trains a GPT from scratch and has an AI agent iteratively modify train.py to minimize val_bpb (bits per byte).
📊 Dataset Stats
Metric
Value
Total documents
19,730
Total text
98.9 MB (~103M characters)… See the full description on the dataset page: https://huggingface.co/datasets/CollinL/perovskite-solar-cell-efficiency-autoresearch.solar-panel-orientationdetails_upstage__SOLAR-10.7B-Instruct-v1.0
Dataset Card for Evaluation run of upstage/SOLAR-10.7B-Instruct-v1.0
Dataset automatically created during the evaluation run of model upstage/SOLAR-10.7B-Instruct-v1.0.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_upstage__SOLAR-10.7B-Instruct-v1.0.solar-rl
SolarChain-Eval RL Benchmark Data
This dataset contains the release bundle for SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets.
The benchmark evaluates autonomous economic governors in decentralized solar-energy markets. It combines city-level photovoltaic generation, peer-to-peer demand, market liquidity, token-burn dynamics, physics-constraint checks, baseline policies, trained RL policies, no-physics ablations… See the full description on the dataset page: https://huggingface.co/datasets/global-nomad-nexus/solar-rl.matbind-datasolana-clawd-instruct
Solana Clawd Instruct
A curated instruction-tuning dataset for fine-tuning models into Solana-native Clawd agents with strong Solana, DeFi, ZK, and constitutional-alignment coverage.
What it teaches
Check every domain your dataset covers:
Solana mechanics (PDAs, accounts, instructions, rent, compute budgets, Token-2022)
DeFi primitives (AMMs, CLMMs, perpetuals, bonding curves, Jupiter, Phoenix)
Memecoin risk analysis (rug detection, holder concentration… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-instruct.AnneStokesThis is a dataset that is based off of the works of Anne Stokes, it's made using Pirsus Artstation which is trained off of SD 1.5 ...the images have been cropped, touched up, and resized to SD 1.5's base resolutions...512x768 and 768x512.
...you should be able to use kohya or dreambooth to train a lora using this.
india-solar-benchmark-dataset
India Solar Benchmark Dataset
A large-scale benchmark dataset for solar irradiance forecasting and renewable energy research built from NASA POWER meteorological observations across 50 major Indian cities.
The benchmark contains 10 years of hourly observations (2016–2025) and is distributed as two complementary datasets:
india_multicity_raw.parquet – cleaned and standardized observations after preprocessing, intended for custom feature engineering and research.… See the full description on the dataset page: https://huggingface.co/datasets/Narendersingh007/india-solar-benchmark-dataset.solana-clawd-repo-corpus
Solana Clawd Core AI Instruct
Instruction-tuning dataset derived from the local core-ai source tree and the
existing Solana Clawd AI training corpus.
Contents
Total examples: 441
Existing ai-training SFT examples: 0
Core AI source chunk examples: 0
Core AI knowledge JSONL examples: 0
Format
Each row is a chat conversation in OpenAI/Hugging Face messages schema:
{"messages": [{"role": "system", "content": "..."}, {"role": "user", "content": "..."}… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-repo-corpus.solar-panel-fault-datasetdetails_upstage__SOLAR-10.7B-v1.0
Dataset Card for Evaluation run of upstage/SOLAR-10.7B-v1.0
Dataset automatically created during the evaluation run of model upstage/SOLAR-10.7B-v1.0.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_upstage__SOLAR-10.7B-v1.0.solar-panels-IGN-bdorthoA YOLO large model was trained with this dataset on IGN BdOrtho imagery.
The resulted detection has very good results which can be see on the MapRoulette project https://maproulette.org/browse/projects/62887.
The trained model is there : https://huggingface.co/Cyrille37/solar-panels-IGN-bdortho
solana-clawd-nvidia-trading-factory-instruct
Solana Clawd NVIDIA Trading Factory Instruct
Specialized SFT data for a Solana-native NVIDIA algorithmic trading factory.
It teaches data ingestion, GPU feature engineering, alpha research, cuML KDE
scenario generation, cuFOLIO/cuOpt Mean-CVaR optimization, paper execution
policy, risk controls, backtesting, monitoring, and Clawd governance.
Format
Each row uses OpenAI-style messages plus metadata:
{"messages": [{"role": "system", "content": "..."}, {"role":… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-nvidia-trading-factory-instruct.
