embedded
Datasets
All datasets matching “embedded”FineWeb2-embedded
FineWeb2-embedded
Dataset summary
FineWeb2-embedded is an extension of the FineWeb2 dataset, annotated with document-level XLM-RoBERTa embeddings for 20 languages, making the dataset useful for a variety of tasks, including document clustering, filtering, and other multilingual research.
Since XLM-RoBERTa has a sequence length limit of 512 tokens, each document's embeddings are obtained by mean-pooling 512 token chunks of the XLM-RoBERTa output. Therefore, longer texts… See the full description on the dataset page: https://huggingface.co/datasets/epfml/FineWeb2-embedded.wdc-common-crawl-embedded-jsonldopenwebtext-t5refinedweb-embedded_prototypeContaining 768-dimensional embedding vectors derived from the "content" column of the RefinedWeb dataset.
The embeddings were generated using the E5-Base-4k model with a context length of 1024, employing scaled dot-product attention.
Original dataset:
This dataset bases as derivative work on RefinedWeb, an English web dataset created by the TII (Technology Innovation Institute).
Attribution is given to the TII as original authors of the RefinedWeb dataset as per Section 4 of RefinedWeb's Open… See the full description on the dataset page: https://huggingface.co/datasets/Marcus2112/refinedweb-embedded_prototype.xsum_validation_t5Wikipedia-TR-2023-Embedded-Dump
Wikipedia-TR-2023-Embedded-Dump
Türkçe Vikipedi (tr.wikipedia.org) makaleleri, retrieval / RAG kullanım
senaryoları için parçalanmış (chunk) ve embedding'lenmiş hâliyle. Her
makale bir parent (ana) chunk'a (tüm makale metni, bağlam
genişletmek için) ve birden fazla child (alt) chunk'a (her biri
kendi embedding'ine sahip küçük pasajlar) bölünmüştür.
İçerik
Makale
348.751
Embedding'li child chunk
1.308.623
Parent chunk (embeddingsiz)
348.751… See the full description on the dataset page: https://huggingface.co/datasets/SalihHub/Wikipedia-TR-2023-Embedded-Dump.
