datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
postgresql-llm
postgresql-llm
A pure PostgreSQL dataset for training and evaluating LLMs on PostgreSQL SQL and PL/pgSQL. Every row is a (question, schema, SQL) triplet with rich metadata for filtering and analysis.
Dataset Summary
postgresql-llm is a pure PostgreSQL dataset: SQL and PL/pgSQL only, with metadata for difficulty, category, and source.
Metric
Value
Total rows
211,539
PostgreSQL-specific rows
11,998 (5.7%)
Schema fill rate
82.2%
Explanation fill rate
17.8%… See the full description on the dataset page: https://huggingface.co/datasets/neurondb/postgresql-llm.Quant-CoT-Factor-Reasoning-PreviewQuantitative Factor Generation: Chain-of-Thought (CoT) Trajectories
Dataset Description
This is a 100-episode preview of a proprietary Reinforcement Learning from Environment Feedback (RLEF) dataset. It is designed to fine-tune Large Language Models (LLMs) for institutional quantitative finance, specifically systematic factor discovery and vectorized Python execution.
The Architecture
The data captures multi-turn agentic loops where the LLM:
Formulates a cross-sectional equity factor… See the full description on the dataset page: https://huggingface.co/datasets/1Happy-neuron/Quant-CoT-Factor-Reasoning-Preview.python-code-simplified
What This Is
This is a dataset of "simplified" Python Code from https://huggingface.co/datasets/iamtarun/python_code_instructions_18k_alpaca.
Our simplification attempts to remove all comments, and reduce these names/strings to a single letter without impacting its structure/logic.
You should use the "minimized" field. The "original" field is the original output from iamtarun/python_code_instructions_18k_alpaca.
Why It Exists
We used this dataset to generate… See the full description on the dataset page: https://huggingface.co/datasets/neuronpedia-org/python-code-simplified.tha-bmr-qa-datasetneuronspark-phase1-dataneuronspark-sft-data
