datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
syntaxgym_sentencesCode-Syntax-Expanded
Code-Syntax-Expanded
A massive, high-quality synthetic dataset for training LLMs to identify and correct syntax errors across 33 programming languages. Contains 5+ million unique examples (~1.1 GB) with English explanations – no artificial padding, no duplicate rows.
📊 Dataset Overview
Property
Value
Total rows
5,000,000+
File size
~1.1 GB (uncompressed CSV)
Languages
33
Unique templates
160+ error patterns
Format
CSV (4 columns)
License… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/Code-Syntax-Expanded.weather-prediction-prototype-aws
Weather prediction prototype database.
This database was made using data provided by KMI.
This database will only be used to train a prototype.
Dataset Details
Dataset Description
Dataset Sources [optional]
KMI
Dataset Structure
Normalized columns:
timestamp
air_pressure
relative_humidity
precipitation
wind_speed
wind_direction
More information about these columns can be found in the information_10min.txt file.
Code-Syntax-Expanded
Code-Syntax-Expanded
A massive, high-quality synthetic dataset for training LLMs to identify and correct syntax errors across 33 programming languages. Contains 5+ million unique examples (~1.1 GB) with English explanations – no artificial padding, no duplicate rows.
📊 Dataset Overview
Property
Value
Total rows
5,000,000+
File size
~1.1 GB (uncompressed CSV)
Languages
33
Unique templates
160+ error patterns
Format
CSV (4 columns)
License… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Code-Syntax-Expanded.Code-Syntax
Code Syntax Dataset (S)
A large-scale, high‑quality dataset for teaching large language models to identify and correct common syntax errors across 30+ programming languages.Contains 500,000+ unique examples (≈110 MB) with English explanations – no artificial padding.
📊 Dataset Format
The dataset is provided as a single CSV file with the following columns:
Column
Type
Description
wrong_code
string
Code snippet containing a syntax error
correct_code… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Code-Syntax.
