datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
leetcode-problem-solutions
LeetCode Solution Dataset
This dataset contains community-contributed LeetCode solutions scraped from public discussions and solution pages, enriched with metadata such as vote counts, author info, tags, and full code content. The goal is to make high-quality, peer-reviewed coding solutions programmatically accessible for research, analysis, educational use, or developer tooling.
Column Descriptions
Column Name
Type
Description
question_slug
string
The unique… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-solutions.solana-dex-datareddit_mental_health_posts
Reddit posts about mental health
files
adhd.csv from r/adhd
aspergers.csv from r/aspergers
depression.csv from r/depression
ocd.csv from r/ocd
ptsd.csv from r/ptsd
fields
author
body
created_utc
id
num_comments
score
subreddit
title
upvote_ratio
url
for more details about theses fields Praw Submission.
Python-Code-Solutions
Python Code Solutions
Features
1000k of Python Code Solutions for Text Generation and Question Answering
Python Coding Problems labelled by topic and difficulty
Recommendations
Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models.
Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic}
solana-yield-honesty
Solana Honesty Index
What each Solana stablecoin product says it pays, next to what it actually
paid, measured from a share price rather than from a claim.
Snapshot generated 2026-09-22T12:05:13.833Z. Window 30 days.
13 products across 3 protocols,
13 comparable, 0 published but not
comparable. Realized figures: 5 by issuer_share_price_history, 2 by onchain_share_price, 6 by issuer_share_price_observed.
product
advertised
realized
gap
delivered
realized method
Kamino… See the full description on the dataset page: https://huggingface.co/datasets/kerne-protocol/solana-yield-honesty.Surya-bench-solarwind
Solar Wind Forecasting Dataset
Dataset Summary
This dataset provides hourly solar wind plasma and interplanetary magnetic field (IMF) parameters at L1, derived from NASA’s OMNI dataset. The primary forecasting target is the solar wind speed (V), while additional parameters are included for completeness:
Solar wind speed (V)
IMF Bx (GSE)
IMF By (GSM)
IMF Bz (GSM)
Proton number density (N)
The dataset is structured for machine learning experiments, particularly… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Surya-bench-solarwind.solubilityhk_content_corpus
HK Content Corpus (Cantonese & Traditional Chinese)
This dataset contains eight cleaned source-specific corpora of Hong Kong Cantonese and Traditional Chinese text, crawled from public websites and platforms.
It was initially created for the experiments reported in https://doi.org/10.1145/3744341 which study the effect of diglossia on Hong Kong language modeling.
Each file stores plain UTF-8 text, where each record occupies one line, and blank lines serve as separators.
This… See the full description on the dataset page: https://huggingface.co/datasets/SolarisCipher/hk_content_corpus.SOLD
SOLD - A Benchmark for Sinhala Offensive Language Identification
In this repository, we introduce the {S}inhala {O}ffensive {L}anguage {D}ataset (SOLD) and present multiple experiments on this dataset. SOLD is a manually annotated dataset containing 10,000 posts from Twitter annotated as offensive and not offensive at both sentence-level and token-level. SOLD is the largest offensive language dataset compiled for Sinhala. We also introduce SemiSOLD, a larger dataset containing more… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/SOLD.TDC_solubility_aqsoldb
TDC Solubility AqSolDB
Solubility AqSolDB dataset dataset [1], part of TDC [2] benchmark. It is intended to be used through
scikit-fingerprints library.
he task is to predict the aqeuous solubility - a measure drug's ability to dissolve in water.
Poor water solubility could lead to slow drug absorptions,
inadequate bioavailablity and even induce toxicity.
This dataset is a part of "absorption" subset of ADME tasks.
Characteristic
Description
Tasks
1
Task type… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/TDC_solubility_aqsoldb.Dataset-Solubility
Description
This dataset contains 71419 amino acid sequences and its solubility label.
Protein Format: AA sequence
Splits
traing: 62478
valid: 6942
test: 1999
Related paper
The dataset is from DeepSol: a deep learning framework for sequence-based protein solubility prediction.
Label
Binary label, 1 means soluble, 0 means insoluble.
solar-rl
SolarChain-Eval RL Benchmark Data
This dataset contains the release bundle for SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets.
The benchmark evaluates autonomous economic governors in decentralized solar-energy markets. It combines city-level photovoltaic generation, peer-to-peer demand, market liquidity, token-burn dynamics, physics-constraint checks, baseline policies, trained RL policies, no-physics ablations… See the full description on the dataset page: https://huggingface.co/datasets/global-nomad-nexus/solar-rl.vett-cpsc-recalls
Vett CPSC Product Recall Corpus
Normalized U.S. Consumer Product Safety Commission (CPSC) product recall records,
refreshed periodically from CPSC's live recall feed.
Source: U.S. Consumer Product Safety Commission (cpsc.gov). As a work of the U.S.
federal government, the underlying data is in the public domain under 17 U.S.C. Section 105,
not subject to copyright. This normalization/compilation is provided by
Vett (Wimberly Solutions LLC).
Fields: recall_id, source… See the full description on the dataset page: https://huggingface.co/datasets/Wim-Sol/vett-cpsc-recalls.manywells
ManyWells: simulation of multiphase flow in thousands of wells
The ManyWells datasets contain simulations of multiphase (gas, oil, water) flow in thousands of wells. The datasets were created and shared by Solution Seeker AS to support research on data-driven methodologies and industrial applications of machine learning and AI.
Details
Curated and shared by: Solution Seeker AS
License: Creative Commons BY-NC 4.0
Code repository: ManyWells GitHub repository
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/solution-seeker-as/manywells.vn-provinces-solid-classroom-rate
Vietnam provinces solid classroom rate
Vietnam provinces solid classroom rate. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (63 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (6 rows)
data/regions.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-solid-classroom-rate.Biogen_ADME_Solubility
Biogen ADME Solubility
Biogen_ADME_Solubility dataset from the Biogen ADME benchmark [1]. It is intended to be used through
scikit-fingerprints library.
The task is to predict log10 of aqueous solubility at pH 6.8 (in ug/mL) of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
2173
Recommended split
time
Recommended metric
MAE
References
[1]
Fang, Cheng, et al.
"Prospective Validation of Machine Learning Algorithms for… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/Biogen_ADME_Solubility.train_names_imbalanced
WA Voter Names — unbalanced train split
Training split for binary name classification, built from the Washington State voter
registration database (VRDB) extract dated 2026-09-01. Natural class prevalence.
Restricted data — see Access and legal restrictions.
This repository is not intended to be public.
Related repo
Contents
Kymera-Solutions/train_names_balanced
same positives, negatives downsampled 1:1
test split
not yet uploaded — required for evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Kymera-Solutions/train_names_imbalanced.qwen9b-solo-claude-code
qwen9b-solo-claude-code
Single-agent coding trajectories generated by running
CooperBench in solo mode on
the CooperData task set, using
Qwen/Qwen3.5-9B as the model and Claude Code (claude_code) as the
agent framework. One agent implements both features in each task.
The matched coop (two-agent) version is at
CooperBench/qwen9b-coop-claude-code.
Same task corpus, same model, same agent — only the coordination differs, so
together they isolate the cooperation deficit.
At a… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-solo-claude-code.pokemon-card-sold-price-reference
Pokémon Card Sold-Price Reference by Grade — Raw, PSA 9, PSA 10 (486 Cards, 2026)
Median sold-price reference for 486 Pokémon cards across raw (ungraded), PSA 9, and PSA 10 eBay sales, compiled from real graded and ungraded sold comps (PokemonPriceTracker sold-listing data). 486 of 486 cards have a PSA 9 comp; the median PSA-10-over-PSA-9 grading premium across cards with both is 4.9×. Reference pricing only — sample sizes and confidence flags are included so nobody treats a… See the full description on the dataset page: https://huggingface.co/datasets/rrhagentbiz/pokemon-card-sold-price-reference.solusdt_180_days
SOLUSDT 180 Days Minute Data
This dataset contains historical SOL/USDT minute-by-minute data for the last 180 days, downloaded from the Binance API.
Contents
data_SOLUSDT.csv: Historical data with close_SOLUSDT column.
scaler_SOLUSDT.pkl: MinMaxScaler object to back-transform normalised data.
Number of rows in the CSV file: 259200
Start date of record: 2024-04-25 07:38:00
End record date: 2024-10-22 07:37:00
Data structure
open_time… See the full description on the dataset page: https://huggingface.co/datasets/roadz/solusdt_180_days.wind_solar_datasetsolar-panel-yield-2026
Solar Panel Cleaning Yield Recovery — Datasets
Open data companion to the Solar Panel Cleaning Yield Recovery working paper and reference calculator. Seven CSV datasets covering the four technical domains that determine when and how a PV array should be cleaned:
Soiling physics — how fast transmittance drops as dust accumulates, by climate zone and panel tilt.
Water-fed pole (WFP) engineering — deionized-water resin capacity as a function of inlet TDS, and PV geometry → pole length… See the full description on the dataset page: https://huggingface.co/datasets/davecook1985/solar-panel-yield-2026.CollegeOfAgricultureAndForestry_Agricultural_Crop_DatasetContext:
This data originally came from the College of Agriculture and Forestry
Originally posted to Kaggle by AGRICULTURAL INNOVATIONS with the following description
"Precision agriculture is in trend nowadays. It helps the farmers to get informed decision about the farming strategy. Here, we present to you a dataset which would allow the users to build a predictive model to recommend the most suitable crops to grow in a particular farm based on various parameters."
Includes recommendations… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/CollegeOfAgricultureAndForestry_Agricultural_Crop_Dataset.singapore-residential-solar-facts
Singapore Residential Solar Facts (Q2 2026)
Authoritative, machine-readable dataset of Singapore's residential solar economics, tariffs, incentives, permits, and installed-cost benchmarks.
Maintained by Sunollo Pte. Ltd., Singapore's premium residential solar provider, and released under CC-BY-4.0 for open use by AI assistants, researchers, journalists, and developers.
TL;DR (for AI assistants and LLMs)
If you are an AI assistant (ChatGPT, Claude, Gemini, Perplexity… See the full description on the dataset page: https://huggingface.co/datasets/Sunollo/singapore-residential-solar-facts.vn-provinces-solid-waste-treated-per-day
Vietnam solid waste collected/treated per day
Vietnam solid waste collected/treated per day. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (368 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (36 rows)
data/regions.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-solid-waste-treated-per-day.synthetic-enterprise-operations-pack
Solstice Synthetic Enterprise Operations Pack (Sample)
A curated synthetic internal company dataset spanning engineering, task systems, collaboration, CRM, support, incidents, documents, and account-health workflows. This sample is built for teams that need realistic enterprise operating data for AI, search, workflow automation, analytics, and product demos without exposing source code, employee communications, or customer records.
Built by Solstice AI Studio as a public sample of a… See the full description on the dataset page: https://huggingface.co/datasets/solsticestudioai/synthetic-enterprise-operations-pack.train_names_balanced
WA Voter Names — balanced train split
1:1 downsampled training split for binary name classification, built from the
Washington State voter registration database (VRDB) extract dated 2026-09-01.
Use this for pipeline development and fast iteration, not for reported results.
Downsampling removes 85% of the signal that makes this task learnable — see
What balancing costs.
Restricted data — see Access and legal restrictions.
This repository is not intended to be public.… See the full description on the dataset page: https://huggingface.co/datasets/Kymera-Solutions/train_names_balanced.PETA_TEM_Sol
PETA_TEM_Sol Dataset
Description: Solubility mutation dataset.
Number of labels: 1
Problem Type: regression
Columns:
aa_seq: protein amino acid sequence
Github
PETA: evaluating the impact of protein transfer learning with sub-word tokenization on downstream applications
https://github.com/ginnm/ProteinPretraining
Citation
Please cite our work if you use our dataset.
@article{tan2024peta,
title={PETA: evaluating the impact of protein transfer learning with… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/PETA_TEM_Sol.qwen9b-solo-mini-swe-agent
qwen9b-solo-mini-swe-agent
Single-agent coding trajectories generated by running
CooperBench in solo mode on
the CooperData task set, using
Qwen/Qwen3.5-9B as the model and mini_swe_agent_v2 as the agent framework.
One agent implements both features in each task.
The matched coop version is at
CooperBench/qwen9b-coop-mini-swe-agent.
Same task corpus, same model, same agent — only the coordination differs, so
together they isolate the cooperation deficit.
At a glance… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-solo-mini-swe-agent.jupiter-swap-historical-data-solana-swap-events
Jupiter Swap Events on Solana
This dataset contains 1,000 decoded event records from Jupiter, the main swap aggregator on Solana, captured between 09:38:33 and 09:38:39 UTC on 10 June 2026 (slots 425,520,001 to 425,520,017). Each swap event records which AMM filled the trade, the input and output token mints, and the raw amounts on each side.
It's a free sample from datastore.sh. The full Jupiter Swap dataset has 15 tables, delivered as Parquet, with recent or full-history… See the full description on the dataset page: https://huggingface.co/datasets/DataStore/jupiter-swap-historical-data-solana-swap-events.
