CoolFace
Datasetpublic

fromziro/jetoncount_corpus

JetonCount's Corpus This is the corpus used to train JetonCount. JSONL Format { "source_file": "token_stats\\HuggingFaceFW_fineweb-edu00000.jsonl", "dataset_dir": "HuggingFaceFW/fineweb-edu", "index": 4020, "chars": 2178, "words": 335, "avg_chars_per_word": 5.504478, "longest_word_chars": 33, "punctuation_ratio": 0.037649, "symbol_ratio": 0.00551, "tokens": 664, "vocab_size": 2560, "tokenizer_dir": "fromziro/Er-Tiny-1.3M" }… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/jetoncount_corpus.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
1likes198downloads

fromziro/jetoncount_corpus · main · files are served by the source, never re-hosted here