CoolFace
Datasetpublic

ashvardanian/StringWars

StringKilla - Small Datasets for String Algorithms Benchmarking The goal of this dataset is to provide a fairly diverse set of strings to evalute the performance of various string-processing algorithms in StringZilla and beyond. English Texts English Leipzig Corpora Collection 124 MB uncompressed 1'000'000 lines of ASCII 8'388'608 tokens of mean length 5 The dataset was originally pulled from Princeton's website: wget --no-clobber -O leipzig1M.txt… See the full description on the dataset page: https://huggingface.co/datasets/ashvardanian/StringWars.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
1likes136downloads
settings

This repository belongs to ashvardanian on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameStringWars
visibilitypublic
licenceapache-2.0
gatedno
ownerashvardanian
Account settings
ashvardanian/StringWars · CoolFace