CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cpllab /syntaxgym_sentencestabular1K<n<10K1 likes107 downloads4y agoHugging Face02Corpus-NZ /Code-Syntax-Expanded Code-Syntax-Expanded A massive, high-quality synthetic dataset for training LLMs to identify and correct syntax errors across 33 programming languages. Contains 5+ million unique examples (~1.1 GB) with English explanations – no artificial padding, no duplicate rows. 📊 Dataset Overview Property Value Total rows 5,000,000+ File size ~1.1 GB (uncompressed CSV) Languages 33 Unique templates 160+ error patterns Format CSV (4 columns) License… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/Code-Syntax-Expanded.text10M<n<100M0 likes60 downloads28d agoHugging Face03syntaxnoob /weather-prediction-prototype-aws Weather prediction prototype database. This database was made using data provided by KMI. This database will only be used to train a prototype. Dataset Details Dataset Description Dataset Sources [optional] KMI Dataset Structure Normalized columns: timestamp air_pressure relative_humidity precipitation wind_speed wind_direction More information about these columns can be found in the information_10min.txt file. tabular100K<n<1M1 likes46 downloads3y agoHugging Face04Gugu8 /Code-Syntax-Expanded Code-Syntax-Expanded A massive, high-quality synthetic dataset for training LLMs to identify and correct syntax errors across 33 programming languages. Contains 5+ million unique examples (~1.1 GB) with English explanations – no artificial padding, no duplicate rows. 📊 Dataset Overview Property Value Total rows 5,000,000+ File size ~1.1 GB (uncompressed CSV) Languages 33 Unique templates 160+ error patterns Format CSV (4 columns) License… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Code-Syntax-Expanded.text10M<n<100M0 likes27 downloads2mo agoHugging Face05Gugu8 /Code-Syntax Code Syntax Dataset (S) A large-scale, high‑quality dataset for teaching large language models to identify and correct common syntax errors across 30+ programming languages.Contains 500,000+ unique examples (≈110 MB) with English explanations – no artificial padding. 📊 Dataset Format The dataset is provided as a single CSV file with the following columns: Column Type Description wrong_code string Code snippet containing a syntax error correct_code… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Code-Syntax.text1M<n<10M0 likes20 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.