datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
solana-memecoin-calls
Solana memecoin calls — a public record with the misses left in
8,161 pump.fun token calls, each with the market cap we called it at, the peak it reached
afterwards, and the exact second it was posted publicly. The whole file is hashed and the hash is
anchored in a Bitcoin block, so no row can be added, edited or back-dated after the fact.
Every trading channel publishes its winners. This is the same feed with the losers still in it —
about six calls in ten never double, and… See the full description on the dataset page: https://huggingface.co/datasets/Smurfetc/solana-memecoin-calls.Solace-270K-Golden-131K-SFT
Solace-270K-Golden-131K-SFT
Official 270,000 Golden Distillation Corpus for 131K Native Context Post-Training
Executive Summary
Solstice-AI/Solace-270K-Golden-131K-SFT is the curated, high-purity post-training corpus created by Solstice-AI, extracted and balanced from the landmark 12.59M-conversation Solstice-AI/Solace-1.0-Omni foundation.
Designed specifically for 131,072 Token (131K Token) native context post-training, this dataset contains zero… See the full description on the dataset page: https://huggingface.co/datasets/Solstice-AI/Solace-270K-Golden-131K-SFT.solana-pairs-history
Dataset Card for Solana Pairs History
This dataset card provides an overview of the "Solana Pairs Price History", a collection of historical data related to Solana liquidity pairs. It is intended for use in research and development of financial models, data analysis, and machine learning applications.
Dataset Details
Dataset Description
The dataset contains historical trading data for Solana pairs, with each pair represented as a separate JSONL file. The… See the full description on the dataset page: https://huggingface.co/datasets/horenresearch/solana-pairs-history.solar-flare-hmivideo-datasplitsThis dataset is intended to be used for training/testing solar flare forecasting models.
It contains various data splits (in json format) of SDO/HMI magnetogram images compiled by Boucheron, L.E., et al., 2023, Sci Data 10, 825, https://doi.org/10.1038/s41597-023-02628-8 and arranged in 16 video frames with different frame cadence (32 min, 76 min).
Splits "train_72min" ("train_36min"), "val_72min" ("val_36min"), "test_72min" ("val_36min") corresponds to the original data splits provided by… See the full description on the dataset page: https://huggingface.co/datasets/inaf-oact-ai/solar-flare-hmivideo-datasplits.upstage__SOLAR-10.7B-v1.0-details
Dataset Card for Evaluation run of upstage/SOLAR-10.7B-v1.0
Dataset automatically created during the evaluation run of model upstage/SOLAR-10.7B-v1.0
The dataset is composed of 83 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 23 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/upstage__SOLAR-10.7B-v1.0-details.solar-inverter-panel-battery-compatibility
Solar inverter–battery compatibility
Canonical, always-current version: https://referencesource.org/solar-inverter-panel-battery-compatibility/
Machine-readable: https://referencesource.org/solar-inverter-panel-battery-compatibility/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-12
Stale after: 2027-02-08 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 101
Which lithium batteries have been… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/solar-inverter-panel-battery-compatibility.Solana-Vanguard-Challenge
Solana Vanguard Challenge Dataset
Overview
The Solana Vanguard Challenge dataset is an official benchmark designed to evaluate and train AI models on the full spectrum of Solana ecosystem expertise and smart contract programming skills. With 1,000 carefully curated questions, this dataset spans foundational concepts, advanced on-chain development in Rust (including the Anchor framework), and sophisticated client-side integration with TypeScript. It is intended for… See the full description on the dataset page: https://huggingface.co/datasets/Bifrost-AI/Solana-Vanguard-Challenge.solarVQA
SolarVQA: A Benchmark for Visual Question Answering in Photovoltaic Defect Inspection
Dataset Summary
SolarVQA is a structured Visual Question Answering dataset built from
expert-annotated electroluminescence (EL) images of silicon solar cells.
It contains 130,712 QA pairs across 16,339 images spanning eight
complementary question types designed to probe defect existence, counting,
type identification, severity, localisation, co-occurrence, and spatial
distribution.… See the full description on the dataset page: https://huggingface.co/datasets/masked-token/solarVQA.sola-code-1-data
sola-code-1-data
Training + eval data for Sola-Code-1 (Solana Anchor code reviewer).
Family Sola by Krene — Sola general, Sola-Code coder. Anchor is the framework tag, never the name.
Contents
train.jsonl — 2000 rows (400 × account-validation, pda, cpi-safety, reentrancy, economic)
eval.jsonl — 200 held-out rows (40 × 5), NEVER trained on
SHA256SUMS — eval hash (leak check)
Row shape
{"instruction": "Review this Anchor program for vulnerabilities"… See the full description on the dataset page: https://huggingface.co/datasets/YukoNikumo/sola-code-1-data.solana-mev-bundle-telemetry
Solana Jito MEV Bundle Telemetry & Execution Traces
High-frequency telemetry dataset of sub-slot block engine bundle auctions, arbitrage executions, and tip distribution metrics on Solana mainnet. Generated from production searcher nodes operating across Raydium CPMM/CLMM, Orca Whirlpools, Meteora DLMM, and Pump.fun bonding curves via Jito Block Engine relays.
Official Implementations & Production Bots
This open research telemetry dataset is maintained in… See the full description on the dataset page: https://huggingface.co/datasets/Mevboters/solana-mev-bundle-telemetry.upstage__solar-pro-preview-instruct-details
Dataset Card for Evaluation run of upstage/solar-pro-preview-instruct
Dataset automatically created during the evaluation run of model upstage/solar-pro-preview-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/upstage__solar-pro-preview-instruct-details.upstage__SOLAR-10.7B-Instruct-v1.0-details
Dataset Card for Evaluation run of upstage/SOLAR-10.7B-Instruct-v1.0
Dataset automatically created during the evaluation run of model upstage/SOLAR-10.7B-Instruct-v1.0
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/upstage__SOLAR-10.7B-Instruct-v1.0-details.solar-flare-hmi-datasplitsThis dataset is intended to be used for training/testing solar flare forecasting models. It contains various data splits (in json format) of SDO/HMI magnetogram images compiled by
Boucheron, L.E., et al., 2023, Sci Data 10, 825, https://doi.org/10.1038/s41597-023-02628-8.
Splits "train", "val", "test" corresponds to the original data splits provided by Boucheron et al., while the other splits are created by downsampling the No-Flare and C flare class to obtain
more balanced splits and… See the full description on the dataset page: https://huggingface.co/datasets/inaf-oact-ai/solar-flare-hmi-datasplits.NousResearch__Yarn-Solar-10b-32k-details
Dataset Card for Evaluation run of NousResearch/Yarn-Solar-10b-32k
Dataset automatically created during the evaluation run of model NousResearch/Yarn-Solar-10b-32k
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Yarn-Solar-10b-32k-details.NousResearch__Yarn-Solar-10b-64k-details
Dataset Card for Evaluation run of NousResearch/Yarn-Solar-10b-64k
Dataset automatically created during the evaluation run of model NousResearch/Yarn-Solar-10b-64k
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Yarn-Solar-10b-64k-details.solana-tx-foundation-unifiedNousResearch__Nous-Hermes-2-SOLAR-10.7B-details
Dataset Card for Evaluation run of NousResearch/Nous-Hermes-2-SOLAR-10.7B
Dataset automatically created during the evaluation run of model NousResearch/Nous-Hermes-2-SOLAR-10.7B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Nous-Hermes-2-SOLAR-10.7B-details.alibi-interrogation-dataset
Alibi Interrogation Dataset
1,736 synthetic multi-turn interrogation conversations for fine-tuning LLMs to play murder suspects. Built for the "Alibi" interrogation game concept.
Format
ShareGPT-style JSONL. Each line is a conversation:
{"conversations": [{"from": "system", "value": "Character card..."}, {"from": "human", "value": "Detective question"}, {"from": "gpt", "value": "Suspect response with *physical cues*"}, ...]}
Compatible with Unsloth, Axolotl, and any… See the full description on the dataset page: https://huggingface.co/datasets/solarkyle/alibi-interrogation-dataset.Solana_150freewheelin__free-solar-evo-v0.13-details
Dataset Card for Evaluation run of freewheelin/free-solar-evo-v0.13
Dataset automatically created during the evaluation run of model freewheelin/free-solar-evo-v0.13
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/freewheelin__free-solar-evo-v0.13-details.solar-flare-goes-datasetThis dataset is intended to be used for training/testing solar flare forecasting models. It contains various data splits (in json format) of
GOES XRS time series (1 min-cadence) for two variables:
L2 flux/bkg ratio
Flare binary history (0=no flare, 1=flare)
Splits labelled as "_24h" correspond to a time series length of 24h, while those labelled as "_12h" to a length of 12h.
Biology_German_DHBWSolana_vulnerability_audit_dataset_V2
Uploaded Dataset
Name: Solana_vulnerability_audit_dataset_V2
Organization: Armur
Project: Solana Smart Contract Audit
License: apache-2.0
Language: en
SKZF6Jq9y0K67iCmSolana_1000VAGOsolutions__SauerkrautLM-SOLAR-Instruct-details
Dataset Card for Evaluation run of VAGOsolutions/SauerkrautLM-SOLAR-Instruct
Dataset automatically created during the evaluation run of model VAGOsolutions/SauerkrautLM-SOLAR-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/VAGOsolutions__SauerkrautLM-SOLAR-Instruct-details.Weyaxi__SauerkrautLM-UNA-SOLAR-Instruct-details
Dataset Card for Evaluation run of Weyaxi/SauerkrautLM-UNA-SOLAR-Instruct
Dataset automatically created during the evaluation run of model Weyaxi/SauerkrautLM-UNA-SOLAR-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Weyaxi__SauerkrautLM-UNA-SOLAR-Instruct-details.zKS0qsmMY9161NhYV3Eez69swvOO9tGKLmU12QfppBWgkdfO
