datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
postgresql-llm
postgresql-llm
A pure PostgreSQL dataset for training and evaluating LLMs on PostgreSQL SQL and PL/pgSQL. Every row is a (question, schema, SQL) triplet with rich metadata for filtering and analysis.
Dataset Summary
postgresql-llm is a pure PostgreSQL dataset: SQL and PL/pgSQL only, with metadata for difficulty, category, and source.
Metric
Value
Total rows
211,539
PostgreSQL-specific rows
11,998 (5.7%)
Schema fill rate
82.2%
Explanation fill rate
17.8%… See the full description on the dataset page: https://huggingface.co/datasets/neurondb/postgresql-llm.NeuronSpark-V1
NeuronSpark-V1 Pretraining Dataset
Bilingual (English + Chinese) pretraining corpus for NeuronSpark, a bio-inspired Spiking Neural Network language model.
Dataset Summary
Metric
Value
Total documents
17,174,734
Estimated tokens
~14.5B
Languages
English (55%), Chinese (42%), Bilingual Math (3%)
Format
Parquet (35 shards, ~39 GB)
Columns
text (string), source (string)
Sources & Composition
Source
Documents
Ratio
Est. Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Brain2nd/NeuronSpark-V1.NeuronSpark-Pretrain-v3
NeuronSpark-Pretrain-v3
Bilingual pretraining corpus for NeuronSpark v3, a bio-inspired Spiking Neural
Network language model with selective PLIF neurons and dynamic per-token compute
budget (PonderNet-v3).
Composition
Metric
Value
Total documents
18.2 M
Estimated tokens
~20 B
Format
37 Parquet shards (~1 GB each, zstd)
Schema
text: string, source: string
Languages
EN 55.6%, ZH 28.1%, code 16.3%
Deduplication
All source sampling is weighted so each… See the full description on the dataset page: https://huggingface.co/datasets/Brain2nd/NeuronSpark-Pretrain-v3.neuron_33Symptom2Diseaseneuronpedia-sae-concepts
Neuronpedia SAE Concepts
Complete extraction of all individual concepts from every Sparse Autoencoder (SAE) released on Neuronpedia, plus all public features from Anthropic's Towards Monosemanticity (2023) and Scaling Monosemanticity (2024) papers.
Quick Start
from datasets import load_dataset
# Full Neuronpedia dataset (77M rows, streaming recommended)
ds = load_dataset("hbe/neuronpedia-sae-concepts", split="train", streaming=True)
# Unique concepts with essential… See the full description on the dataset page: https://huggingface.co/datasets/hbe/neuronpedia-sae-concepts.NeuronSpark-SFT-Mixpythia-12b-neuron-dataset-examples
pythia-12b-neuron-dataset-examples
This dataset contains the top 64 highest activating dataset examples for each
MLP neuron in Pythia-12b. The dataset examples are all 16 tokens long. See
https://confirmlabs.org/posts/dreaming.html for details.
Columns:
layer: the layer of the neuron
neuron: the index of the neuron
rank: the rank of the example
activation: the activation of the neuron on the example
position: the token position for which the neuron is maximally activated.
text: the… See the full description on the dataset page: https://huggingface.co/datasets/Confirm-Labs/pythia-12b-neuron-dataset-examples.neuronovo-utc-data-glue-mnlillama8b-layer15-meta-neurons
Llama8B Meta-Neurons
This repository contains meta-neuron data accompanying the paper Learning a Generative Meta-Model of LLM Activations.
Project page: https://generative-latent-prior.github.io
Code: https://github.com/g-luo/generative_latent_prior
Quick Start
With this data, you can browse the 98304 meta-neurons of the Llama-3.1-8B GLP (glp-llama8b-d6, Layer 15).
Meta-neurons are the post-SwiGLU activations of the GLP's MLP blocks. For each meta-neuron… See the full description on the dataset page: https://huggingface.co/datasets/generative-latent-prior/llama8b-layer15-meta-neurons.amazonneuronovo-utc-data-goemotionsuaspeech_train_castedgsm8k-uz
GSM8K-UZ
Uzbek (Latin script) translation of openai/gsm8k
(main config), produced with nvidia/gemma-4-31B-it-NVFP4.
Split
Rows
Source rows
Retained
train
7,417
7,473
99.25%
test
1,308
1,319
99.17%
Columns
Column
Description
question
The problem in Uzbek
answer
The step-by-step solution in Uzbek, with the original <<...>> calculator annotations and the final #### N line preserved
question_en
The original English question
answer_en… See the full description on the dataset page: https://huggingface.co/datasets/NeuronUz/gsm8k-uz.neuronovo-utc-persent-docSTELAR-topo_vision_reasoning_SFT_50k
Stellar-Neuron/STELAR-topo_vision_reasoning_SFT_50k
[Paper] [HF Collection] [Project Page]
The dataset was released as part of STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision. STELAR is a more accurate, faster and greener intelligent system for Vision Language Reasoning.
Contact: chenli4@andrew.cmu.edu
Dataset Summary
This dataset was created by STELAR TopoAug from two base datasets: Math-V and VLM_S2H. Each question includes… See the full description on the dataset page: https://huggingface.co/datasets/Stellar-Neuron/STELAR-topo_vision_reasoning_SFT_50k.neuronovo-utc-data-glue-colaSTELAR-topo_vision_reasoning_preference_123k
Stellar-Neuron/STELAR-topo_vision_reasoning_preference_123k
[Paper] [HF Collection] [Project Page]
The dataset was released as part of STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision. STELAR is a more accurate, faster and greener intelligent system for Vision Language Reasoning.
Contact: chenli4@andrew.cmu.edu
Dataset Summary
This dataset was created by STELAR TopoAug from two base datasets: Math-V and VLM_S2H.
Each question includes… See the full description on the dataset page: https://huggingface.co/datasets/Stellar-Neuron/STELAR-topo_vision_reasoning_preference_123k.torgo_full_dataset_with_idneuronovo-utc-unhealthy-conversationsneuronovo-utc-measuring-hate-speechargilla_preferences_personalized_filteredneuronovo-utc-tweeteval-sentimentrlhf_synthetic_generalizedsma-gse108094-motor-neuron-rnaseq
GSE108094 SMA Motor-Neuron RNA-seq
This repository contains processed RNA-sequencing results and GEO/SRA metadata for human SMA and control iPSC-derived motor neurons.
The experiment has eight libraries from four biological cell lines: two SMA and two control lines, each with two sequencing replicates. The repository contains the differential-expression output, alternative-splicing output, series matrix, MINiML family file, SOFT family file, and SRA run inventory. The… See the full description on the dataset page: https://huggingface.co/datasets/YannisTevissen/sma-gse108094-motor-neuron-rnaseq.neuronovo-utc-tweeteval-emotionsSTELAR-topo_vision_reasoning_100k
Stellar-Neuron/STELAR-topo_vision_reasoning_100k
[Paper] [HF Collection] [Project Page]
The dataset was released as part of STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision. STELAR is a more accurate, faster and greener intelligent system for Vision Language Reasoning.
Contact: chenli4@andrew.cmu.edu
Dataset Summary
This dataset was created by STELAR TopoAug from two base datasets: Math-V and VLM_S2H. Each question includes responses… See the full description on the dataset page: https://huggingface.co/datasets/Stellar-Neuron/STELAR-topo_vision_reasoning_100k.Airlines_Reviews_Neuronettorgo_full_datasetuaspeech_test
