akhilhsingh/homeo-dataset
๐ท FineWeb 15 trillion tokens of the finest data the ๐ web has to offer What is it? The ๐ท FineWeb dataset consists of more than 15T tokens of cleaned and deduplicated english web data from CommonCrawl. The data processing pipeline is optimized for LLM performance and ran on the ๐ญ datatrove library, our large scale data processing library. ๐ท FineWeb was originally meant to be a fully open replication of ๐ฆ RefinedWeb, with a release of the fullโฆ See the full description on the dataset page: https://huggingface.co/datasets/akhilhsingh/homeo-dataset.
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Create README.md
Upload homeo.pdf
initial commit
