CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01eaglewatch /Korean_Wikipedia_Dataset_for_GPT2_August_2022 Dataset Card for korean_wikipedia_dataset_for_GPT2 Dataset Description Entire Korean language Wikipedia data for GPT-2 training as of August 1st, 2022. email: oscar.eaglewatch@gmail.com Dataset Summary This is to make a pre-trained GPT-2 Korean model Languages Korean Dataset Structure Data Instances Train wikipedia article count: 334420 validation wikipedia article count: 83605 Data Fields 'text' Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/eaglewatch/Korean_Wikipedia_Dataset_for_GPT2_August_2022.textquestion-answering100K<n<1M6 likes95 downloads2y agoHugging Face02WithinUsAI /gpt2_to_gpt5.5_distilled_25k GPT-2 to GPT-5.5 Advanced Reasoning Distillation (25k) Dataset Description 25,000 unique, high-quality instruction-response pairs designed for knowledge distillation and supervised fine-tuning. The dataset elevates GPT-2 Medium toward GPT-5.5-level performance on complex reasoning tasks. Core goal: Transfer frontier reasoning capabilities (multi-step CoT, cross-domain synthesis, edge-case analysis, novel insights) from a hypothetical GPT-5.5 teacher into smaller… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/gpt2_to_gpt5.5_distilled_25k.texttext-generation10K<n<100K1 likes22 downloads4mo agoHugging Face03nobodyiam /gpt2-dataset Hello Datasets This is the dataset used to fine tune fine-tuned-gpt2. textquestion-answeringn<1K0 likes9 downloads3y agoHugging Face04kurtos-ai /gsm-gpt2-rff gsm-gpt2-rff A reproducible, standardized mathematical reasoning dataset constructed from five public sources, with controllable data selection fractions (γ) using a Generalized Linear Model with Random Fourier Feature for selection. Overview This dataset is an aggregate of five widely used mathematical reasoning datasets: deepmind/aqua_rat (raw, train split) openai/gsm8k (main, train split) allenai/math_qa (train split) meta-math/MetaMathQA (train split)… See the full description on the dataset page: https://huggingface.co/datasets/kurtos-ai/gsm-gpt2-rff.textquestion-answering100K<n<1M0 likes7 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.