karpathy/tiny_shakespeare
40,000 lines of Shakespeare from a variety of Shakespeare's plays. Featured in Andrej Karpathy's blog post 'The Unreasonable Effectiveness of Recurrent Neural Networks': http://karpathy.github.io/2015/05/21/rnn-effectiveness/. To use for e.g. character modelling: ``` d = datasets.load_dataset(name='tiny_shakespeare')['train'] d = d.map(lambda x: datasets.Value('strings').unicode_split(x['text'], 'UTF-8')) # train split includes vocabulary for other splits vocabulary = sorted(set(next(iter(d)).numpy())) d = d.map(lambda x: {'cur_char': x[:-1], 'next_char': x[1:]}) d = d.unbatch() seq_len = 100 batch_size = 2 d = d.batch(seq_len) d = d.batch(batch_size) ```
Delete legacy JSON metadata (#3)
Convert dataset sizes from base 2 to base 10 in the dataset card (#1)
add dataset_info in dataset metadata
remove dummmy data
Remove a copy-paste sentence in dataset cards (#4281)
Update files from the datasets library (from 1.18.0)
Update files from the datasets library (from 1.7.0)
Update files from the datasets library (from 1.6.0)
Update files from the datasets library (from 1.4.0)
Update files from the datasets library (from 1.3.0)
Update files from the datasets library (from 1.0.0)
