CoolFace
Datasetpublic

Trelis/tiny-shakespeare

Data source Downloaded via Andrej Karpathy's nanogpt repo from this link Data Format The entire dataset is split into train (90%) and test (10%). All rows are at most 1024 tokens, using the Llama 2 tokenizer. All rows are split cleanly so that sentences are whole and unbroken.

sourceHugging Faceupdated 3y agoView on Hugging Face
11likes14kdownloads
Dataset Card

Data source

Downloaded via Andrej Karpathy's nanogpt repo from this link

Data Format

  • The entire dataset is split into train (90%) and test (10%).
  • All rows are at most 1024 tokens, using the Llama 2 tokenizer.
  • All rows are split cleanly so that sentences are whole and unbroken.