CoolFace
Datasetpublic

alphass123/TinyStories

Dataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in the following paper: https://arxiv.org/abs/2305.07759. The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M. Additional resources: tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/alphass123/TinyStories.

sourceHugging Facecdla-sharing-1.0updated 6mo agoView on Hugging Face
0likes15downloads
../
filetrain-00000-of-00004-2d5a1467fff1081b.parquet237.2 MBdownload
filetrain-00001-of-00004-5852b56a2bd28fd9.parquet236.7 MBdownload
filetrain-00002-of-00004-a26307300439e943.parquet234.5 MBdownload
filetrain-00003-of-00004-d243063613e5a057.parquet236.5 MBdownload
filevalidation-00000-of-00001-869c898b519ad725.parquet9.5 MBdownload

alphass123/TinyStories · main · files are served by the source, never re-hosted here