agentlans/finewebedu-sentences
Fineweb-edu Sentences Description: A dataset of sentences collected from the web. The dataset was created by splitting the text into individual sentences using the spaCy package, then removing duplicates and filtering for complete sentences in a semi-automated process. Source: HuggingFaceFW/fineweb-edu Size: About 700,000 English language sentences. Each sentence is 512 tokens long or less as assessed using the BERT tokenizer. Annotations: The source field contains the URL of… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/finewebedu-sentences.
025
Upload train.csv.gz
Update README.md
Upload train.csv.gz
initial commit
