datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Korean_Wikipedia_Dataset_for_GPT2_August_2022
Dataset Card for korean_wikipedia_dataset_for_GPT2
Dataset Description
Entire Korean language Wikipedia data for GPT-2 training as of August 1st, 2022.
email: oscar.eaglewatch@gmail.com
Dataset Summary
This is to make a pre-trained GPT-2 Korean model
Languages
Korean
Dataset Structure
Data Instances
Train wikipedia article count: 334420
validation wikipedia article count: 83605
Data Fields
'text'
Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/eaglewatch/Korean_Wikipedia_Dataset_for_GPT2_August_2022.gpt2_to_gpt5.5_distilled_25k
GPT-2 to GPT-5.5 Advanced Reasoning Distillation (25k)
Dataset Description
25,000 unique, high-quality instruction-response pairs designed for knowledge distillation and supervised fine-tuning. The dataset elevates GPT-2 Medium toward GPT-5.5-level performance on complex reasoning tasks.
Core goal: Transfer frontier reasoning capabilities (multi-step CoT, cross-domain synthesis, edge-case analysis, novel insights) from a hypothetical GPT-5.5 teacher into smaller… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/gpt2_to_gpt5.5_distilled_25k.gpt2-dataset
Hello Datasets
This is the dataset used to fine tune fine-tuned-gpt2.
gsm-gpt2-rff
gsm-gpt2-rff
A reproducible, standardized mathematical reasoning dataset constructed from five public sources, with controllable data selection fractions (γ) using a Generalized Linear Model with Random Fourier Feature for selection.
Overview
This dataset is an aggregate of five widely used mathematical reasoning datasets:
deepmind/aqua_rat (raw, train split)
openai/gsm8k (main, train split)
allenai/math_qa (train split)
meta-math/MetaMathQA (train split)… See the full description on the dataset page: https://huggingface.co/datasets/kurtos-ai/gsm-gpt2-rff.
