CoolFace
Datasetpublic

dark-xet/test-public-dataset

Dataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in the following paper: https://arxiv.org/abs/2305.07759. The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M. Additional resources: tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/dark-xet/test-public-dataset.

sourceHugging Facecdla-sharing-1.0updated 2y agoView on Hugging Face
0likes113downloads
2 commits on main
2aa4c532y ago

Upload folder using huggingface_hub

jsulz
1be91c72y ago

initial commit

jsburner