CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01agentlans /finewebedu-sentences Fineweb-edu Sentences Description: A dataset of sentences collected from the web. The dataset was created by splitting the text into individual sentences using the spaCy package, then removing duplicates and filtering for complete sentences in a semi-automated process. Source: HuggingFaceFW/fineweb-edu Size: About 700,000 English language sentences. Each sentence is 512 tokens long or less as assessed using the BERT tokenizer. Annotations: The source field contains the URL of each… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/finewebedu-sentences.text100K<n<1M0 likes27 downloads2y agoHugging Face02cemig-ceia /fineweb-edu-gemini-annotations-portuguese-regressiontabular1K<n<10K0 likes9 downloads1y agoHugging Face03irfanfadhullah /FineWeb-Edu-25Ktabular10K<n<100K0 likes9 downloads1y agoHugging Face04alexandermorgan /FineWeb-Edu_10B_sample_2_column_word_countsThis dataset contains is a 2-column csv file representation of the FineWeb-Edu 10B sample dataset which is made available under the Open Data Commons Attribution License (ODC-By). All the texts in the ~27GB of parquet files were split according to the regex below using the Python regex package (not re). The text chunks from these splits were counted to make the str: int mapping of the csv file. The result is a greater than 100X reduction in file size (27GB -> 240MB) making it easy to fit this… See the full description on the dataset page: https://huggingface.co/datasets/alexandermorgan/FineWeb-Edu_10B_sample_2_column_word_counts.text10M<n<100M0 likes8 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.