CoolFace
Datasetpublic

agentlans/finewebedu-sentences

Fineweb-edu Sentences Description: A dataset of sentences collected from the web. The dataset was created by splitting the text into individual sentences using the spaCy package, then removing duplicates and filtering for complete sentences in a semi-automated process. Source: HuggingFaceFW/fineweb-edu Size: About 700,000 English language sentences. Each sentence is 512 tokens long or less as assessed using the BERT tokenizer. Annotations: The source field contains the URL of… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/finewebedu-sentences.

sourceHugging Faceodc-byupdated 2y agoView on Hugging Face
0likes25downloads
Dataset Card

Fineweb-edu Sentences

Description: A dataset of sentences collected from the web. The dataset was created by splitting the text into individual sentences using the spaCy package, then removing duplicates and filtering for complete sentences in a semi-automated process.

Source: HuggingFaceFW/fineweb-edu

Size: About 700,000 English language sentences. Each sentence is 512 tokens long or less as assessed using the BERT tokenizer.

Annotations: The source field contains the URL of each sentence from FineWeb-Edu.

License: Open Data Commons License Attribution (same as HuggingFaceFW/fineweb-edu)