fromziro/jetoncount_corpus
JetonCount's Corpus This is the corpus used to train JetonCount. JSONL Format { "source_file": "token_stats\\HuggingFaceFW_fineweb-edu00000.jsonl", "dataset_dir": "HuggingFaceFW/fineweb-edu", "index": 4020, "chars": 2178, "words": 335, "avg_chars_per_word": 5.504478, "longest_word_chars": 33, "punctuation_ratio": 0.037649, "symbol_ratio": 0.00551, "tokens": 664, "vocab_size": 2560, "tokenizer_dir": "fromziro/Er-Tiny-1.3M" }… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/jetoncount_corpus.
1198
