CoolFace
Datasetpublic

enjalot/fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5

FineWeb-edu 10BT Sample embedded with nomic-text-v1.5 The FineWeb-edu 10BT sample was first chunked into 500 tokens (using bert-base-uncased) with 10% overlap resulting in 25 million rows and 10.5BT. The chunks were then embedded using nomic-text-v1.5. Dataset Details Dataset Sources Repository: https://github.com/enjalot/fineweb-modal Uses Direct Use The dataset was embedded with the clustering: prefix, so the main… See the full description on the dataset page: https://huggingface.co/datasets/enjalot/fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
5likes647downloads
4 commits on main
f4125b82y ago

Update README.md

enjalot
97922d92y ago

Upload dataset (part 00001-of-00002)

enjalot
ca3d71a2y ago

Upload dataset (part 00000-of-00002)

enjalot
e03e99a2y ago

initial commit

enjalot