datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
postgresql-llm
postgresql-llm
A pure PostgreSQL dataset for training and evaluating LLMs on PostgreSQL SQL and PL/pgSQL. Every row is a (question, schema, SQL) triplet with rich metadata for filtering and analysis.
Dataset Summary
postgresql-llm is a pure PostgreSQL dataset: SQL and PL/pgSQL only, with metadata for difficulty, category, and source.
Metric
Value
Total rows
211,539
PostgreSQL-specific rows
11,998 (5.7%)
Schema fill rate
82.2%
Explanation fill rate
17.8%… See the full description on the dataset page: https://huggingface.co/datasets/neurondb/postgresql-llm.neuronpedia-sae-concepts
Neuronpedia SAE Concepts
Complete extraction of all individual concepts from every Sparse Autoencoder (SAE) released on Neuronpedia, plus all public features from Anthropic's Towards Monosemanticity (2023) and Scaling Monosemanticity (2024) papers.
Quick Start
from datasets import load_dataset
# Full Neuronpedia dataset (77M rows, streaming recommended)
ds = load_dataset("hbe/neuronpedia-sae-concepts", split="train", streaming=True)
# Unique concepts with essential… See the full description on the dataset page: https://huggingface.co/datasets/hbe/neuronpedia-sae-concepts.gaia-dr3-oa-neuron-xp-spectra
Gaia DR3 OA neuron XP spectra
This table contains the prototype BP/RP spectrum attached to each neuron in the 30 × 30 self-organising map produced by Gaia's Apsis Outlier Analysis module. Measured from the served table, every (neuron_id, xp_spectrum_prototype_wavelength) pair is unique across its 78,300 rows. Each of the 900 neurons has the same ordered grid of 87 wavelengths from 379.0935 to 1034.2460 nm. Neuron-level statistics and other attributes belong to Gaia's separate… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gaia-dr3-oa-neuron-xp-spectra.pythia-12b-neuron-dataset-examples
pythia-12b-neuron-dataset-examples
This dataset contains the top 64 highest activating dataset examples for each
MLP neuron in Pythia-12b. The dataset examples are all 16 tokens long. See
https://confirmlabs.org/posts/dreaming.html for details.
Columns:
layer: the layer of the neuron
neuron: the index of the neuron
rank: the rank of the example
activation: the activation of the neuron on the example
position: the token position for which the neuron is maximally activated.
text: the… See the full description on the dataset page: https://huggingface.co/datasets/Confirm-Labs/pythia-12b-neuron-dataset-examples.neuronovo-utc-data-glue-mnlillama8b-layer15-meta-neurons
Llama8B Meta-Neurons
This repository contains meta-neuron data accompanying the paper Learning a Generative Meta-Model of LLM Activations.
Project page: https://generative-latent-prior.github.io
Code: https://github.com/g-luo/generative_latent_prior
Quick Start
With this data, you can browse the 98304 meta-neurons of the Llama-3.1-8B GLP (glp-llama8b-d6, Layer 15).
Meta-neurons are the post-SwiGLU activations of the GLP's MLP blocks. For each meta-neuron… See the full description on the dataset page: https://huggingface.co/datasets/generative-latent-prior/llama8b-layer15-meta-neurons.coffee_sales_dataneuronovo-utc-data-goemotionsneuronovo-utc-persent-docneuronovo-utc-data-glue-colaLlama-3.1-8B_neuron-activationsneuronovo-utc-unhealthy-conversationsneuronovo-utc-measuring-hate-speechLlama-3.2-3B_neuron-activationsargilla_preferences_personalized_filteredneuronovo-utc-tweeteval-sentimentrlhf_synthetic_generalizedsma-gse108094-motor-neuron-rnaseq
GSE108094 SMA Motor-Neuron RNA-seq
This repository contains processed RNA-sequencing results and GEO/SRA metadata for human SMA and control iPSC-derived motor neurons.
The experiment has eight libraries from four biological cell lines: two SMA and two control lines, each with two sequencing replicates. The repository contains the differential-expression output, alternative-splicing output, series matrix, MINiML family file, SOFT family file, and SRA run inventory. The… See the full description on the dataset page: https://huggingface.co/datasets/YannisTevissen/sma-gse108094-motor-neuron-rnaseq.Airlines_Reviews_Neuronetneuronovo-utc-tweeteval-emotionsQuant-CoT-Factor-Reasoning-PreviewQuantitative Factor Generation: Chain-of-Thought (CoT) Trajectories
Dataset Description
This is a 100-episode preview of a proprietary Reinforcement Learning from Environment Feedback (RLEF) dataset. It is designed to fine-tune Large Language Models (LLMs) for institutional quantitative finance, specifically systematic factor discovery and vectorized Python execution.
The Architecture
The data captures multi-turn agentic loops where the LLM:
Formulates a cross-sectional equity factor… See the full description on the dataset page: https://huggingface.co/datasets/1Happy-neuron/Quant-CoT-Factor-Reasoning-Preview.neuronovo-utc-hate-speech18-sentencesOLMo-7B-0424-hf_neuron-activationsThis dataset contains activation data of neurons in OLMo-7B-0424.
(We define a neuron as a hidden dimension in a MLP sublayer.)
To create the dataset, the model was run on 20M tokens from Dolma (in the same collection, we also release the Dolma subset, which we call dolma-small).
Dataset Description
Each row corresponds to a neuron, identified by the columns "layer" and "neuron".
(We use zero-based indexing).
The other columns are as follows:
The first two elements of the name… See the full description on the dataset page: https://huggingface.co/datasets/sjgerstner/OLMo-7B-0424-hf_neuron-activations.amortized-neuron-pruning-effects
Amortized Neuron-Ablation Effects (for pruning)
Exact mean-ablation effects of individual MLP neurons in transformer LMs, paired
with cheap forward/backward signals, for the task of amortized causal
intervention-effect prediction at neuron granularity (pruning).
The intervention unit is a single MLP neuron (blocks.{l}.mlp.hook_post channel).
For every (prompt, neuron) we record the exact effect of replacing that neuron's
last-token activation with its dataset mean, m(x… See the full description on the dataset page: https://huggingface.co/datasets/almogtavor/amortized-neuron-pruning-effects.sfd_housing_prices_november_2024
Index designations for "author_type_id":
Realtor – 0
Homeowner – 1
Index designations for "location_id":
Astrakhan – 1
Volgograd – 2
Krasnodar – 3
Rostov-on-Don – 4
Maykop – 5
Elista – 6
Index designations for "district_id":
Astrakhan:
Kirovsky - 11
Leninsky - 12
Sovietsky - 13
Trusovsky - 14
Central - 15
Volgograd:
Voroshilovsky - 21
Dzerzhinsky - 22
Kirovsky - 23
Krasnoarmeysky - 24
Krasnooktyabrsky - 25
Sovietsky - 26
Traktorozavodsky - 27
Central -… See the full description on the dataset page: https://huggingface.co/datasets/neuronetties/sfd_housing_prices_november_2024.neuronalTTR_targetsTranscriptionActivators_molecularFeaturesThe neuronalTTR_targetsTranscriptionActivators_molecularFeatures dataset is a part from the study “Comparative analysis of computational approaches for predicting Transthyretin (TTR) transcription activators and human dopamine D1 receptor antagonists”
https://doi.org/10.48550/arXiv.2506.01137
A total of 3,041 unique small molecule samples are included in this dataset. The samples are classified by their TTR transcription activity, resulting in 1,093 activators and 1,948 non-activators. This… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/neuronalTTR_targetsTranscriptionActivators_molecularFeatures.neuronalTTR_targetsTranscriptionActivators_C13NMRThe neuronalTTR_targetsTranscriptionActivators_13CNMR dataset is a part from the study “Comparative analysis of computational approaches for predicting Transthyretin (TTR) transcription activators and human dopamine D1 receptor antagonists”
https://doi.org/10.48550/arXiv.2506.01137
A total of 3,041 unique small molecule samples are included in this dataset. The samples are classified by their TTR transcription activity, resulting in 1,093 activators and 1,948 non-activators. This information… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/neuronalTTR_targetsTranscriptionActivators_C13NMR.qwen-h-neurons-datasetcar_fault_typeneuronalTTR_targetsTranscriptionActivators_CID_SIDThe neuronalTTR_targetsTranscriptionActivators_CID_SID dataset is a part from the study “Comparative analysis of computational approaches for predicting Transthyretin (TTR) transcription activators and human dopamine D1 receptor antagonists”
https://doi.org/10.48550/arXiv.2506.01137
A total of 3,174 unique small molecule samples are included in this dataset. The samples are classified by their TTR transcription activity, resulting in 2,020 activators and 1,154 non-activators. This… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/neuronalTTR_targetsTranscriptionActivators_CID_SID.
