tinyllm
Datasets
All datasets matching “tinyllm”tiny-llm-synthetic-qa
Tiny-LLM: Synthetic Question-Answering Dataset
Dataset Description
This dataset was created for the fine-tuning stage of the Tiny-LLM Project, a project focused on training and evaluating compact language models from scratch.
It contains 706,727 high-quality, synthetic multi-turn Question-Answering (Q&A) conversations in English, generated using the Gemini API. The dataset was designed to teach small models instruction-following capabilities across a diverse range of… See the full description on the dataset page: https://huggingface.co/datasets/Gabriel8/tiny-llm-synthetic-qa.aime-1983-2023-trajectoriesgpqa-extended-trajectoriestiny-llm
tiny-llm retained training evidence
1782 seed/horizon records across 594 configurations: cosine-to-zero and WSD
schedules, synchronous and four- and eight-worker decentralized training, at 20 to 160
global tokens per parameter, for 20.4M-parameter models on C4. Snapshot 2026-09-17.
This mirrors doc/data/current-training/ in
WangZesen/tiny-llm. The website built from it is
at https://wangzesen.github.io/tiny-llm/.
Layout
Retained artifacts are grouped into gzipped… See the full description on the dataset page: https://huggingface.co/datasets/zesen-kth/tiny-llm.tinyllm-datagpqa-main-trajectories
