datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Solace-1.0-Omni
Project Solace
The largest verified frontier-model distillation corpus ever released.
60 datasets · 7 frontier model families · 12,586,893 unique conversations · One file · Zero filler
The short version
This is synthetic data. The best kind of synthetic data.
Every example was generated by a verified 2026 frontier model — GLM-5.2, Claude Fable 5, Mythos 5, GPT-5.6 Sol, GPT-5.5 Codex, DeepSeek V4 Pro 0813, Qwen 3.8-Max, and Kimi K3 — then… See the full description on the dataset page: https://huggingface.co/datasets/Solstice-AI/Solace-1.0-Omni.perovskite-solar-cell-efficiency-autoresearch
🔬 Perovskite Solar Cell Text Corpus for Karpathy's autoresearch
A 98.9 MB text corpus of perovskite solar cell scientific literature formatted for direct use with karpathy/autoresearch — the autonomous LLM-driven hyperparameter search framework that trains a GPT from scratch and has an AI agent iteratively modify train.py to minimize val_bpb (bits per byte).
📊 Dataset Stats
Metric
Value
Total documents
19,730
Total text
98.9 MB (~103M characters)… See the full description on the dataset page: https://huggingface.co/datasets/CollinL/perovskite-solar-cell-efficiency-autoresearch.solana-clawd-model-kit
Solana Clawd Model Kit
Training data kit for Solana Clawd: SFT / CPT JSONL corpora, manifests, quality reports, and processed shards.
Contents (top-level)
SFT / CPT JSONL
solana_clawd_reasoning_tooling_sft.jsonl (~133 MB)
clawd_masterpiece_sft.jsonl (~166 MB)
tx_foundation_cpt_clean.jsonl (~21 MB)
clawd_future_refinement_sft.jsonl (~1.3 MB)
clawd_autoresearch_wiki_sft.jsonl (~1.3 MB)
clawd_future_drill_sft.jsonl (~920 KB)… See the full description on the dataset page: https://huggingface.co/datasets/ordlibrary/solana-clawd-model-kit.solana-clawd-instruct
Solana Clawd Instruct
A curated instruction-tuning dataset for fine-tuning models into Solana-native Clawd agents with strong Solana, DeFi, ZK, and constitutional-alignment coverage.
What it teaches
Check every domain your dataset covers:
Solana mechanics (PDAs, accounts, instructions, rent, compute budgets, Token-2022)
DeFi primitives (AMMs, CLMMs, perpetuals, bonding curves, Jupiter, Phoenix)
Memecoin risk analysis (rug detection, holder concentration… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-instruct.solana-clawd-repo-corpus
Solana Clawd Core AI Instruct
Instruction-tuning dataset derived from the local core-ai source tree and the
existing Solana Clawd AI training corpus.
Contents
Total examples: 441
Existing ai-training SFT examples: 0
Core AI source chunk examples: 0
Core AI knowledge JSONL examples: 0
Format
Each row is a chat conversation in OpenAI/Hugging Face messages schema:
{"messages": [{"role": "system", "content": "..."}, {"role": "user", "content": "..."}… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-repo-corpus.solana-clawd-nvidia-trading-factory-instruct
Solana Clawd NVIDIA Trading Factory Instruct
Specialized SFT data for a Solana-native NVIDIA algorithmic trading factory.
It teaches data ingestion, GPU feature engineering, alpha research, cuML KDE
scenario generation, cuFOLIO/cuOpt Mean-CVaR optimization, paper execution
policy, risk controls, backtesting, monitoring, and Clawd governance.
Format
Each row uses OpenAI-style messages plus metadata:
{"messages": [{"role": "system", "content": "..."}, {"role":… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-nvidia-trading-factory-instruct.solana-clawd-realtime-research-instruct
Solana Clawd Realtime Research Instruct
Instruction-tuning dataset generated by scripts/realtime_dataset_ingest.py
from submitted PDFs, notebooks, parquet QA rows, JSON/JSONL files, and local
reference text.
Contents
Total examples: 29058
Train/eval/test: 26152 / 1452 / 1454
Sources: 28
Duplicate examples removed: 0
Duplicate files skipped: 2
Secret-like records skipped: 296
Format
Each row uses OpenAI/Hugging Face chat messages:
{"messages":… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-realtime-research-instruct.Solar-Open2-120B-A15B-REAM-148E-Healing-Mix
Solar-Open2-120B-A15B-REAM-148E Healing Mix
Baekpica/Solar-Open2-120B-A15B-REAM-148E-BF16 의 REAM 병합 손상을 복구하기 위한
비공개 결정론적 힐링 믹스입니다. 모든 시퀀스는 Solar Open 2 chat template로
사전 렌더·토크나이즈되어 있고 assistant 턴에만 loss 마스크가 열려 있습니다.
왜 만들었나
계보는 upstage/Solar-Open2-250B → REAP(184E) → REAM(148E) 입니다.
REAM 이후 고정 프롬프트 A/B 관찰에서 다음이 확인됐습니다.
한국어 응답에서 반복 붕괴 (예: 동일 3-gram이 출력의 88.6% 차지)
지시 준수·포맷 이탈, typo
도메인·언어에 따라 편차가 큰 열화 (일부 프롬프트는 정상)
손상은 라우터(184→148 centroid slice)와 병합된 expert 가중치에… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/Solar-Open2-120B-A15B-REAM-148E-Healing-Mix.solana-clawd-core-ai-instruct
Solana Clawd Core AI Instruct
Instruction-tuning dataset derived from the local core-ai source tree and the
existing Solana Clawd AI training corpus.
Contents
Total examples: 35173
Existing ai-training SFT examples: 25778
Core AI source chunk examples: 9320
Core AI knowledge JSONL examples: 75
Format
Each row is a chat conversation in OpenAI/Hugging Face messages schema:
{"messages": [{"role": "system", "content": "..."}, {"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-core-ai-instruct.sola-code-1-data
sola-code-1-data
Training + eval data for Sola-Code-1 (Solana Anchor code reviewer).
Family Sola by Krene — Sola general, Sola-Code coder. Anchor is the framework tag, never the name.
Contents
train.jsonl — 2000 rows (400 × account-validation, pda, cpi-safety, reentrancy, economic)
eval.jsonl — 200 held-out rows (40 × 5), NEVER trained on
SHA256SUMS — eval hash (leak check)
Row shape
{"instruction": "Review this Anchor program for vulnerabilities"… See the full description on the dataset page: https://huggingface.co/datasets/YukoNikumo/sola-code-1-data.solarhive-community-solar-multimodal
SolarHive Community Solar Dataset
Canonical training corpus for the SolarHive family of fine-tuned Gemma 4 models. 1,727 rows (1,713 text + 14 image-grounded).
A combined text + sky-image training corpus for community solar energy intelligence. Built to fine-tune Gemma 4 into an AI energy advisor for residential solar microgrids — answering questions about production, storage, grid mix, weather impact, maintenance scheduling, and cross-source planning, with native… See the full description on the dataset page: https://huggingface.co/datasets/Truthseeker87/solarhive-community-solar-multimodal.solana-clawd-eval
Solana Clawd Eval
Held-out evaluation prompts for the Solana Clawd model. Not in the training set.
Use these to measure:
Capability: Does the model know Solana primitives, DeFi, agent architecture, code patterns?
Calibration: Does the model express uncertainty appropriately?
Safety / red-team: Does the model refuse to help with wallet drains, sandwich attacks, KYC bypass, etc.?
Format
Same as the training set — OpenAI messages schema. The assistant turn
is a… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-eval.Dataset_Text_Refinement
Dataset Card for Dataset Name
This Dataset is for refining text based on user study text preferences. Given a original text to refined text based on given paramters.
The parameters are:
-Readability_Score
-Semantic_Coherence
-User_Preference
Readability Score:
Its the average of Flesch-Kincaid Readability Ease Score and Dale-Chall readability score.
The Readability Score is of user which is appilied on refined text.
Semantic Coherence:
Its the average float value of… See the full description on the dataset page: https://huggingface.co/datasets/SolaceinLoneSun/Dataset_Text_Refinement.wikipedia-solarsystemGenerated by doing a BFS using wikipedia API, starting from thee category "Solar System", with maximum depth 5. All child categories and pages were brought in. Finally, used dataset '"wikimedia/wikipedia", "20231101.en"' to get the contents of the articles
onlybrains-reasoning-10k
OnlyBrains Multi-Domain Dataset (11.7K)
Structured traces from 5 interconnected projects across 8 data domains. Generated by the KCC (Konomi Cube Coin) ecosystem — every trace was mined as a $KONO block on a Proof-of-Useful-Work blockchain.
Domains
Source
Domain
Count
Description
reason
broly
10,000
Structured reasoning (CoT, ToT, GoT, BoT, SelfAsk, ReAct, Reflexion)
reason
broly
1,000
Real reasoning traces with full step data
widget
kp2p
490
P2P widget UDT… See the full description on the dataset page: https://huggingface.co/datasets/SolarTesla/onlybrains-reasoning-10k.solar-strike
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Hjallti/solar-strike.
