virtualkevin/standardebooks-jsonl
Standard Ebooks JSONL This dataset is a JSONL conversion of the Hugging Face dataset Nelathan/standardebooks. The source dataset contains full-text public domain books sourced from Standard Ebooks. Dataset Structure The dataset has one split, train, stored as two JSONL shards: data/train-00000-of-00002.jsonl - 617 rows data/train-00001-of-00002.jsonl - 616 rows Each line is a JSON object with the same fields as the source parquet dataset: link: URL of the… See the full description on the dataset page: https://huggingface.co/datasets/virtualkevin/standardebooks-jsonl.
023
Convert Standard Ebooks dataset to JSONL
initial commit
