agentlans/c4-en-tokenized
C4 English Tokenized Samples This dataset contains tokenized English samples from the C4 (Colossal Clean Crawled Corpus) dataset for natural language processing (NLP) tasks. The first 125 000 entries from the en split of allenai/c4 were tokenized using spaCy's en_core_web_sm model. Tokens joined with spaces. Features text: Original text from C4 tokenized: The tokenized and space-joined text num_tokens: Number of tokens after tokenization num_punct_tokens: Number… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/c4-en-tokenized.
08
Upload 2 files
Update README.md
Update README.md
Upload train.jsonl.zst
Update README.md
Upload train.jsonl.zst
initial commit
