CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Josephgflowers /Par-Four-Fineweb-Edu-FortifiedDataset Summary This dataset is a filtered subset of the Fineweb-Edu-Fortified dataset. The primary goal of this subset is to reduce the dataset size to a more manageable volume while maintaining high-quality content. It contains three key fields: score, text, and url, focusing on entries with a score of 4 and above, indicating higher relevance and quality of educational content. This dataset can be used for several fine-tuning and model improvement tasks, including model healing, synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified.text1M<n<10M7 likes128 downloads2y agoHugging Face02Josephgflowers /Par-Four-Fineweb-Edu-Fortified-Chemistry-Physics-Astronomy-Math-ReasonExtracted Chemistry Physics Asronomy Math and Logic portions from the original. Script used for the extraction: https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Chemistry-Physics-Astronomy-Math-Reason/resolve/main/find-science-fine.py tabular100K<n<1M6 likes115 downloads2y agoHugging Face03Josephgflowers /Par-Four-Fineweb-Edu-Fortified-FinanceSubset of https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified Used keyword filtering and scoring to extract data with a finance focus. Creation script: https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Finance/resolve/main/find-fin-fine.py Dataset Description Summary This dataset is a finance-focused filtered subset of the Fineweb-Edu-Fortified dataset. It is designed to extract and prioritize high-quality… See the full description on the dataset page: https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Finance.text100K<n<1M0 likes36 downloads2y agoHugging Face04agentlans /finewebedu-sentences Fineweb-edu Sentences Description: A dataset of sentences collected from the web. The dataset was created by splitting the text into individual sentences using the spaCy package, then removing duplicates and filtering for complete sentences in a semi-automated process. Source: HuggingFaceFW/fineweb-edu Size: About 700,000 English language sentences. Each sentence is 512 tokens long or less as assessed using the BERT tokenizer. Annotations: The source field contains the URL of each… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/finewebedu-sentences.text100K<n<1M0 likes25 downloads2y agoHugging Face05Josephgflowers /Par-Four-Fineweb-Edu-Fortified-MathFiltered for a math focus. Script used to create the dataset: https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Math/resolve/main/find-math-fine.py tabular100K<n<1M1 likes23 downloads2y agoHugging Face06Josephgflowers /Par-Four-Fineweb-Edu-Fortified-LogicFiltered for logic with the following scrip: https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Logic/resolve/main/find-reason-fine.py tabular10K<n<100K0 likes22 downloads2y agoHugging Face07irfanfadhullah /FineWeb-Edu-25Ktabular10K<n<100K0 likes11 downloads1y agoHugging Face08cemig-ceia /fineweb-edu-gemini-annotations-portuguese-regressiontabular1K<n<10K0 likes10 downloads1y agoHugging Face09alexandermorgan /FineWeb-Edu_10B_sample_2_column_word_countsThis dataset contains is a 2-column csv file representation of the FineWeb-Edu 10B sample dataset which is made available under the Open Data Commons Attribution License (ODC-By). All the texts in the ~27GB of parquet files were split according to the regex below using the Python regex package (not re). The text chunks from these splits were counted to make the str: int mapping of the csv file. The result is a greater than 100X reduction in file size (27GB -> 240MB) making it easy to fit this… See the full description on the dataset page: https://huggingface.co/datasets/alexandermorgan/FineWeb-Edu_10B_sample_2_column_word_counts.text10M<n<100M0 likes8 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.