CoolFace
Datasetpublic

agentlans/finewebedu-sentences

Fineweb-edu Sentences Description: A dataset of sentences collected from the web. The dataset was created by splitting the text into individual sentences using the spaCy package, then removing duplicates and filtering for complete sentences in a semi-automated process. Source: HuggingFaceFW/fineweb-edu Size: About 700,000 English language sentences. Each sentence is 512 tokens long or less as assessed using the BERT tokenizer. Annotations: The source field contains the URL of… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/finewebedu-sentences.

sourceHugging Faceodc-byupdated 2y agoView on Hugging Face
0likes25downloads
4 commits on main
124cbdc2y ago

Upload train.csv.gz

agentlans
2159d302y ago

Update README.md

agentlans
8ed90262y ago

Upload train.csv.gz

agentlans
12a94742y ago

initial commit

agentlans