CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01neurondb /postgresql-llm postgresql-llm A pure PostgreSQL dataset for training and evaluating LLMs on PostgreSQL SQL and PL/pgSQL. Every row is a (question, schema, SQL) triplet with rich metadata for filtering and analysis. Dataset Summary postgresql-llm is a pure PostgreSQL dataset: SQL and PL/pgSQL only, with metadata for difficulty, category, and source. Metric Value Total rows 211,539 PostgreSQL-specific rows 11,998 (5.7%) Schema fill rate 82.2% Explanation fill rate 17.8%… See the full description on the dataset page: https://huggingface.co/datasets/neurondb/postgresql-llm.tabulartext-generation100K<n<1M4 likes603 downloads7mo agoHugging Face02Brain2nd /NeuronSpark-V1 NeuronSpark-V1 Pretraining Dataset Bilingual (English + Chinese) pretraining corpus for NeuronSpark, a bio-inspired Spiking Neural Network language model. Dataset Summary Metric Value Total documents 17,174,734 Estimated tokens ~14.5B Languages English (55%), Chinese (42%), Bilingual Math (3%) Format Parquet (35 shards, ~39 GB) Columns text (string), source (string) Sources & Composition Source Documents Ratio Est. Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Brain2nd/NeuronSpark-V1.texttext-generation10M<n<100M0 likes419 downloads6mo agoHugging Face03Brain2nd /NeuronSpark-Pretrain-v3 NeuronSpark-Pretrain-v3 Bilingual pretraining corpus for NeuronSpark v3, a bio-inspired Spiking Neural Network language model with selective PLIF neurons and dynamic per-token compute budget (PonderNet-v3). Composition Metric Value Total documents 18.2 M Estimated tokens ~20 B Format 37 Parquet shards (~1 GB each, zstd) Schema text: string, source: string Languages EN 55.6%, ZH 28.1%, code 16.3% Deduplication All source sampling is weighted so each… See the full description on the dataset page: https://huggingface.co/datasets/Brain2nd/NeuronSpark-Pretrain-v3.texttext-generation10M<n<100M0 likes411 downloads5mo agoHugging Face04NeuronUz /gsm8k-uz GSM8K-UZ Uzbek (Latin script) translation of openai/gsm8k (main config), produced with nvidia/gemma-4-31B-it-NVFP4. Split Rows Source rows Retained train 7,417 7,473 99.25% test 1,308 1,319 99.17% Columns Column Description question The problem in Uzbek answer The step-by-step solution in Uzbek, with the original <<...>> calculator annotations and the final #### N line preserved question_en The original English question answer_en… See the full description on the dataset page: https://huggingface.co/datasets/NeuronUz/gsm8k-uz.texttext-generation1K<n<10K0 likes42 downloads28d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.