CoolFace
Datasetpublic

epfml/FineWeb2-embedded

FineWeb2-embedded Dataset summary FineWeb2-embedded is an extension of the FineWeb2 dataset, annotated with document-level XLM-RoBERTa embeddings for 20 languages, making the dataset useful for a variety of tasks, including document clustering, filtering, and other multilingual research. Since XLM-RoBERTa has a sequence length limit of 512 tokens, each document's embeddings are obtained by mean-pooling 512 token chunks of the XLM-RoBERTa output. Therefore, longer… See the full description on the dataset page: https://huggingface.co/datasets/epfml/FineWeb2-embedded.

sourceHugging Faceodc-byupdated 2y agoView on Hugging Face
6likes26kdownloads
1 commits on main
4631ff22y ago

Super-squash branch 'v1.0.0' using huggingface_hub

vsabolcec