Wikipedia
encoder-e5-small-v2-wikipedia-relationsrelik-reader-deberta-v3-large-re-wikipediaencoder-e5-small-v2-wikipedia-relations-indexbert-base-japanese-v3-ner-wikipedia-datasetword2vec-skipgram-fa-wikipediaword2vec-cbow-fa-wikipediaQwen3.5-9B-Wikipedia-TR-Reasoning-GGUFmpasila_-_Llama-3.2-Finnish-Wikipedia-1B-gguf
wikipedia
Dataset Card for Wikimedia Wikipedia
Dataset Summary
Wikipedia dataset containing cleaned articles of all languages.
The dataset is built from the Wikipedia dumps (https://dumps.wikimedia.org/)
with one subset per language, each containing a single train split.
Each example contains the content of one full Wikipedia article with cleaning to strip
markdown and unwanted sections (references, etc.).
All language subsets have already been processed for recent dump… See the full description on the dataset page: https://huggingface.co/datasets/wikimedia/wikipedia.wikipediaWikipedia dataset containing cleaned articles of all languages.
The datasets are built from the Wikipedia dump
(https://dumps.wikimedia.org/) with one split per language. Each example
contains the content of one full Wikipedia article with cleaning to strip
markdown and unwanted sections (references, etc.).MultiLegalPile_Wikipedia_FilteredA filtered version of the MultiLegalPile dataset, together with wikipedia articles.structured-wikipedia
Dataset Card for Wikimedia Structured Wikipedia
Quick Links
Wikimedia Enterprise
Structured Contents Documentation
Data Dictionary
Wikimedia Attribution Framework
Meta-Wiki Discussion
Dataset Summary
Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API.
This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/wikimedia/structured-wikipedia.wikipedia-2023-11-embed-multilingual-v3
Multilingual Embeddings for Wikipedia in 300+ Languages
This dataset contains the wikimedia/wikipedia dataset dump from 2023-11-01 from Wikipedia in all 300+ languages.
The individual articles have been chunked and embedded with the state-of-the-art multilingual Cohere Embed V3 embedding model. This enables an easy way to semantically search across all of Wikipedia or to use it as a knowledge source for your RAG application. In total is it close to 250M paragraphs / embeddings.
You… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/wikipedia-2023-11-embed-multilingual-v3.wikipedia_culturax_dutch
Filtered CulturaX + Wikipedia for Dutch
This is a combined and filtered version of CulturaX and Wikipedia, only including Dutch. It is intended for the training of LLMs.
Different configs are available based on the number of tokens (see a section below with an overview). This can be useful if you want to know exactly how many tokens you have. Great for using as a streaming dataset, too. Tokens are counted as white-space tokens, so depending on your tokenizer, you'll likely end up… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/wikipedia_culturax_dutch.
