datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Par-Four-Fineweb-Edu-Fortified-Chemistry-Physics-Astronomy-Math-ReasonExtracted Chemistry Physics Asronomy Math and Logic portions from the original.
Script used for the extraction:
https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Chemistry-Physics-Astronomy-Math-Reason/resolve/main/find-science-fine.py
Par-Four-Fineweb-Edu-Fortified-LogicFiltered for logic with the following scrip:
https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Logic/resolve/main/find-reason-fine.py
Par-Four-Fineweb-Edu-Fortified-MathFiltered for a math focus.
Script used to create the dataset:
https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Math/resolve/main/find-math-fine.py
FineWeb-Edu-25Kfineweb-edu-gemini-annotations-portuguese-regressionfineweb_CC-MAIN-2024-18_100k_output_UncovAI_83362
What is it?
As more and more data are generated daily, It becomes important to be able to distinguish between synthetic and human data for model training.
We analyzed the first 100k rows of the Fineweb dataset focusing on the dump CC-MAIN-2024-18 using the UncovAI model for text.
We observed that more than 16% of the data were detected as having been generated by AI by our model. We removed them and obtained a dataset of 83362 lines with a number of token approaching 55 million.… See the full description on the dataset page: https://huggingface.co/datasets/UncovAI/fineweb_CC-MAIN-2024-18_100k_output_UncovAI_83362.fineweb2_ar_15m_sample
