datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
postgresql-llm
postgresql-llm
A pure PostgreSQL dataset for training and evaluating LLMs on PostgreSQL SQL and PL/pgSQL. Every row is a (question, schema, SQL) triplet with rich metadata for filtering and analysis.
Dataset Summary
postgresql-llm is a pure PostgreSQL dataset: SQL and PL/pgSQL only, with metadata for difficulty, category, and source.
Metric
Value
Total rows
211,539
PostgreSQL-specific rows
11,998 (5.7%)
Schema fill rate
82.2%
Explanation fill rate
17.8%… See the full description on the dataset page: https://huggingface.co/datasets/neurondb/postgresql-llm.NeuronSpark-V1
NeuronSpark-V1 Pretraining Dataset
Bilingual (English + Chinese) pretraining corpus for NeuronSpark, a bio-inspired Spiking Neural Network language model.
Dataset Summary
Metric
Value
Total documents
17,174,734
Estimated tokens
~14.5B
Languages
English (55%), Chinese (42%), Bilingual Math (3%)
Format
Parquet (35 shards, ~39 GB)
Columns
text (string), source (string)
Sources & Composition
Source
Documents
Ratio
Est. Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Brain2nd/NeuronSpark-V1.NeuronSpark-Pretrain-v3
NeuronSpark-Pretrain-v3
Bilingual pretraining corpus for NeuronSpark v3, a bio-inspired Spiking Neural
Network language model with selective PLIF neurons and dynamic per-token compute
budget (PonderNet-v3).
Composition
Metric
Value
Total documents
18.2 M
Estimated tokens
~20 B
Format
37 Parquet shards (~1 GB each, zstd)
Schema
text: string, source: string
Languages
EN 55.6%, ZH 28.1%, code 16.3%
Deduplication
All source sampling is weighted so each… See the full description on the dataset page: https://huggingface.co/datasets/Brain2nd/NeuronSpark-Pretrain-v3.gsm8k-uz
GSM8K-UZ
Uzbek (Latin script) translation of openai/gsm8k
(main config), produced with nvidia/gemma-4-31B-it-NVFP4.
Split
Rows
Source rows
Retained
train
7,417
7,473
99.25%
test
1,308
1,319
99.17%
Columns
Column
Description
question
The problem in Uzbek
answer
The step-by-step solution in Uzbek, with the original <<...>> calculator annotations and the final #### N line preserved
question_en
The original English question
answer_en… See the full description on the dataset page: https://huggingface.co/datasets/NeuronUz/gsm8k-uz.
