agentlans/finewebedu-sentences
Fineweb-edu Sentences Description: A dataset of sentences collected from the web. The dataset was created by splitting the text into individual sentences using the spaCy package, then removing duplicates and filtering for complete sentences in a semi-automated process. Source: HuggingFaceFW/fineweb-edu Size: About 700,000 English language sentences. Each sentence is 512 tokens long or less as assessed using the BERT tokenizer. Annotations: The source field contains the URL of… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/finewebedu-sentences.
Fineweb-edu Sentences
Description: A dataset of sentences collected from the web. The dataset was created by splitting the text into individual sentences using the spaCy package, then removing duplicates and filtering for complete sentences in a semi-automated process.
Source: HuggingFaceFW/fineweb-edu
Size: About 700,000 English language sentences. Each sentence is 512 tokens long or less as assessed using the BERT tokenizer.
Annotations: The source field contains the URL of each sentence from FineWeb-Edu.
License: Open Data Commons License Attribution (same as HuggingFaceFW/fineweb-edu)
