datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Par-Four-Fineweb-Edu-FortifiedDataset Summary
This dataset is a filtered subset of the Fineweb-Edu-Fortified dataset. The primary goal of this subset is to reduce the dataset size to a more manageable volume while maintaining high-quality content. It contains three key fields: score, text, and url, focusing on entries with a score of 4 and above, indicating higher relevance and quality of educational content.
This dataset can be used for several fine-tuning and model improvement tasks, including model healing, synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified.Par-Four-Fineweb-Edu-Fortified-Chemistry-Physics-Astronomy-Math-ReasonExtracted Chemistry Physics Asronomy Math and Logic portions from the original.
Script used for the extraction:
https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Chemistry-Physics-Astronomy-Math-Reason/resolve/main/find-science-fine.py
Par-Four-Fineweb-Edu-Fortified-FinanceSubset of https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified
Used keyword filtering and scoring to extract data with a finance focus.
Creation script:
https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Finance/resolve/main/find-fin-fine.py
Dataset Description
Summary
This dataset is a finance-focused filtered subset of the Fineweb-Edu-Fortified dataset. It is designed to extract and prioritize high-quality… See the full description on the dataset page: https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Finance.finewebedu-sentences
Fineweb-edu Sentences
Description:
A dataset of sentences collected from the web.
The dataset was created by splitting the text into individual sentences using the spaCy package,
then removing duplicates and filtering for complete sentences in a semi-automated process.
Source: HuggingFaceFW/fineweb-edu
Size: About 700,000 English language sentences. Each sentence is 512 tokens long or less as assessed using the BERT tokenizer.
Annotations: The source field contains the URL of each… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/finewebedu-sentences.Par-Four-Fineweb-Edu-Fortified-MathFiltered for a math focus.
Script used to create the dataset:
https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Math/resolve/main/find-math-fine.py
Par-Four-Fineweb-Edu-Fortified-LogicFiltered for logic with the following scrip:
https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Logic/resolve/main/find-reason-fine.py
FineWeb-Edu-25Kfineweb-edu-gemini-annotations-portuguese-regressionFineWeb-Edu_10B_sample_2_column_word_countsThis dataset contains is a 2-column csv file representation of the FineWeb-Edu 10B sample dataset which is made available under the Open Data Commons Attribution License (ODC-By). All the texts in the ~27GB of parquet files were split according to the regex below using the Python regex package (not re).
The text chunks from these splits were counted to make the str: int mapping of the csv file. The result is a greater than 100X reduction in file size (27GB -> 240MB) making it easy to fit this… See the full description on the dataset page: https://huggingface.co/datasets/alexandermorgan/FineWeb-Edu_10B_sample_2_column_word_counts.
