CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hotchpotch /wikipedia-ja-20231030 Wikipedia Japanese data (20231030) Source Date: 2023/10/30 Source: https://dumps.wikimedia.org/other/cirrussearch/ License CC BY-SA 4.0 Example WIP tabular1M<n<10M1 likes10k downloads3y agoHugging Face02singletongue /wikipedia-paragraphs wikipedia-paragraphs wikipedia-paragraphs is a dataset generated from Wikipedia, designed for natural language processing (NLP) research. Each entry contains cleaned paragraph text and Wikilink information extracted from a Wikipedia page, along with useful metadata such as categories, templates, and the associated Wikidata QID. Dataset structure Configurations The dataset is organized into multiple configurations, such as enwiki-20260607-v1.2.1.… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/wikipedia-paragraphs.tabular100M<n<1B2 likes9.9k downloads3mo agoHugging Face03NoeFlandre /osm-polygon-wikidata-and-wikipedia OSM Polygon Wikidata + Wikipedia, V2 V2 builds on the V1 Wikidata-only dataset with multilingual Wikipedia and Wikivoyage text. It also retains valid multilingual wikipedia=* references, including polygons without a Wikidata QID. Source code: GitHub repository. Dataset snapshot Metric Value Polygon rows across regional extracts 1,259,424 Unique polygon identities (osm_type, osm_id) 1,188,854 Polygons with successful non-empty text (unique OSM… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-wikidata-and-wikipedia.tabular100M<n<1B0 likes4.6k downloads6h agoHugging Face04maloyan /wikipedia-22-12-en-embeddings-all-MiniLM-L6-v2 Dataset Card for "wikipedia-22-12-en-embeddings-all-MiniLM-L6-v2" More Information needed tabular10M<n<100M4 likes1.6k downloads3y agoHugging Face05Intelligent-Internet /wikipedia_en wikipedia_en This is a curated Wikipedia English dataset for use with the II-Commons project. Dataset Details Dataset Description This dataset comprises a curated Wikipedia English pages. Data sourced directly from the official English Wikipedia database dump. We extract the pages, chunk them into smaller pieces, and embed them using Snowflake/snowflake-arctic-embed-m-v2.0. All vector embeddings are 16-bit half-precision vectors optimized for cosine indexing… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/wikipedia_en.tabularfeature-extraction10M<n<100M2 likes1.3k downloads1y agoHugging Face06singletongue /wikipedia-utils Wikipedia-Utils: Preprocessed Wikipedia Texts for NLP Preprocessed Wikipedia texts generated with the scripts in singletongue/wikipedia-utils repo. For detailed information on how the texts are processed, please refer to the repo. tabular100M<n<1B7 likes1k downloads2y agoHugging Face07vtasca /wikipedia-pageviews Wikipedia Article Pageviews This repository automatically fetches and aggregates the 100 most popular Wikipedia articles by pageviews - creating a dataset that enables tracking trending topics on Wikipedia. It works by polling the WikiMedia API on a daily basis and fetching the top 100 most popular articles from two days ago. The fetcher runs in a scheduled GitHub Actions workflow, which is available here. The dataset begins in the year 2016 and the textual data is presented as… See the full description on the dataset page: https://huggingface.co/datasets/vtasca/wikipedia-pageviews.tabularfeature-extraction100K<n<1M1 likes844 downloads4h agoHugging Face08nthngdy /wikipedia-22-12-concat-split Dataset Card for "wikipedia-22-12-concat-split" More Information needed tabular10M<n<100M0 likes823 downloads3y agoHugging Face09zetavg /zh-tw-wikipedia 台灣正體中文維基百科 (zh-tw Wikipedia) 截至 2023 年 5 月,中文維基百科 2,533,212 篇條目的台灣正體文字內容。每篇條目為一列 (row),包含 HTML 以及 Markdown 兩種格式。 A nearly-complete collection of 2,533,212 Traditional Chinese (zh-tw) Wikipedia pages, gathered between May 1, 2023, and May 7, 2023. Includes both the original HTML format and an auto-converted Markdown version, which has been processed using vinta/pangu.py. 於 2023 年 5 月 1 日至 5 月 7 日間取自維基百科 action=query & prop=extracts API,內容皆與維基百科網站之台灣正體版本一致,沒有繁簡體混雜的問題。 For development… See the full description on the dataset page: https://huggingface.co/datasets/zetavg/zh-tw-wikipedia.tabulartext-generation1M<n<10M30 likes800 downloads3y agoHugging Face10CATIE-AQ /wikipedia_fr_2022French part of Wikipedia embeded with Solon-embeddings-large-0.1. tabular10M<n<100M0 likes799 downloads1y agoHugging Face11Ti-Ma /wikipedia_2017 Dataset Card for Dataset Name This is a Wikipedia dataset correct to "31-12-2017". Dataset Details Dataset Description WikiMedia routinely publishes dumps of Wikipedia, each containing the revision history of articles. We first defined the relevant revision before extracting the article information. Specifically, we select the most recent revision as of December 31st for each year. Consequently, some revisions in our datasets date back several years from the… See the full description on the dataset page: https://huggingface.co/datasets/Ti-Ma/wikipedia_2017.tabular1M<n<10M0 likes738 downloads2y agoHugging Face12Reza2kn /Wikipedia-EN-FA-Accessibility-Bridge Wikipedia EN-FA Accessibility Bridge Current, attributable English and Persian Wikipedia article snapshots for pages created during a 69-day recency window, plus an EN↔FA counterpart index and static accessibility signals. The reproducible full baseline is the official 2026-08-01 Wikimedia dump: 6,289,549 English articles without Persian, 129,821 Persian articles without English, and 22,277,907 namespace-0 pages in the combined parity index. Redirects are retained in the parity… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/Wikipedia-EN-FA-Accessibility-Bridge.tabulartranslation10M<n<100M1 likes675 downloads1mo agoHugging Face13Ti-Ma /wikipedia_2018 Dataset Card for Dataset Name This is a Wikipedia dataset correct to "31-12-2018". Dataset Details Dataset Description WikiMedia routinely publishes dumps of Wikipedia, each containing the revision history of articles. We first defined the relevant revision before extracting the article information. Specifically, we select the most recent revision as of December 31st for each year. Consequently, some revisions in our datasets date back several years from the… See the full description on the dataset page: https://huggingface.co/datasets/Ti-Ma/wikipedia_2018.tabular1M<n<10M0 likes665 downloads2y agoHugging Face14Rubin-Wei /kNN-Targets-wikipedia-mistral Dataset Overview This dataset provides k-nearest neighbor (kNN) target distributions for language modeling. Each token in the Wikipedia corpus is associated with a soft probability distribution over its top-k nearest neighbors in the representation space of a frozen language model. These targets can be used to train MLP Memory. Corresponding Preprocessed Corpus: Rubin-Wei/enwiki-dec2021-preprocessed-mistral Compatible Model: Mistral-7B-v0.3 Paper: MLP Memory: A Retriever-Pretrained… See the full description on the dataset page: https://huggingface.co/datasets/Rubin-Wei/kNN-Targets-wikipedia-mistral.tabular1B<n<10B0 likes640 downloads11mo agoHugging Face15SalihHub /Wikipedia-TR-2023-Embedded-Dump Wikipedia-TR-2023-Embedded-Dump Türkçe Vikipedi (tr.wikipedia.org) makaleleri, retrieval / RAG kullanım senaryoları için parçalanmış (chunk) ve embedding'lenmiş hâliyle. Her makale bir parent (ana) chunk'a (tüm makale metni, bağlam genişletmek için) ve birden fazla child (alt) chunk'a (her biri kendi embedding'ine sahip küçük pasajlar) bölünmüştür. İçerik Makale 348.751 Embedding'li child chunk 1.308.623 Parent chunk (embeddingsiz) 348.751… See the full description on the dataset page: https://huggingface.co/datasets/SalihHub/Wikipedia-TR-2023-Embedded-Dump.tabularfeature-extraction1M<n<10M0 likes590 downloads1mo agoHugging Face16mrm8488 /fineweb-2-spa-5m-wikipedia-2mtabular1M<n<10M0 likes578 downloads1y agoHugging Face17mrm8488 /fineweb-2-spa-3m-wikipedia-2mtabular1M<n<10M0 likes506 downloads1y agoHugging Face18hotchpotch /wikipedia-multilingual-synthetic-ir-query wikipedia-multilingual-synthetic-ir-query This dataset contains multilingual Wikipedia-derived synthetic query-document pairs for information retrieval training. It was created with the query-crafter-multilingual model, which generates search-like queries from Wikipedia text. The current release contains two different retrieval settings: short_doc: pairs of (query, short document) long_doc: pairs of (query, long document) These two subsets are not generated in the same way… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/wikipedia-multilingual-synthetic-ir-query.tabulartext-retrieval10M<n<100M0 likes455 downloads3mo agoHugging Face19ucrelnlp /Multilingual-USAS-Labelled-Silver-Wikipedia Multilingual USAS Silver Labelled Wikipedia Articles Silver-labelled Wikipedia article text for training USAS semantic taggers and Multi-Word Expression (MWE) identifiers, covering 8 Wikipedia language sites. The source text comes from the HuggingFace HuggingFaceFW/finewiki dataset, restricted to articles rated Good (GA) or Featured (FA) — using the article ID list from ucrelnlp/wikipedia-ga-fa-ids — and then sentence split and automatically tagged with USAS semantic tags and… See the full description on the dataset page: https://huggingface.co/datasets/ucrelnlp/Multilingual-USAS-Labelled-Silver-Wikipedia.tabular10K<n<100K0 likes432 downloads4h agoHugging Face20wanng /wikipedia-zh-mnbvc zhwiki-mnbvc 分项目:爬取并处理中文维基百科语料 数据时间:202302-202305 (持续更新) 主项目:MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集 https://github.com/esbatmop/MNBVC 该项目清洗流程主要参考:https://kexue.fm/archives/4176/comment-page-1 并且使用组员开发的去重工具进行数据格式化。 总行数(样本): 10,754,146 一个示例: { "文件名": "cleaned/zhwiki-20230420/folder_0/723712.txt", "是否待查文件": false, "是否重复文件": false, "文件大小": 558, "simhash": 14363740497821204542, "最长段落长度": 142, "段落数": 6, "去重段落数": 6, "低质量段落数": 0, "段落": [ {… See the full description on the dataset page: https://huggingface.co/datasets/wanng/wikipedia-zh-mnbvc.tabulartext-generation1M<n<10M6 likes381 downloads3y agoHugging Face21rcds /wikipedia-for-mask-filling\tabularfill-mask100K<n<1M0 likes306 downloads4y agoHugging Face22seonglae /wikipedia-256-tokenThis is Wikidedia passages dataset for ODQA retriever. Each passages have 256~ tokens splitteed by gpt-4 tokenizer using tiktoken. Token count {'~128': 1415068, '128~256': 1290011, '256~512': 18756476, '512~1024': 667, '1024~2048': 12, '2048~4096': 0, '4096~8192': 0, '8192~16384': 0, '16384~32768': 0, '32768~65536': 0, '65536~128000': 0, '128000~': 0} Text count {'~512': 1556876,'512~1024': 6074975, '1024~2048': 13830329, '2048~4096': 49, '4096~8192': 2, '8192~16384': 3, '16384~32768': 0… See the full description on the dataset page: https://huggingface.co/datasets/seonglae/wikipedia-256-token.tabular10M<n<100M0 likes274 downloads3y agoHugging Face23sirzmkk /zh-tw-wikipedia 台灣正體中文維基百科 (zh-tw Wikipedia) 截至 2023 年 5 月,中文維基百科 2,533,212 篇條目的台灣正體文字內容。每篇條目為一列 (row),包含 HTML 以及 Markdown 兩種格式。 A nearly-complete collection of 2,533,212 Traditional Chinese (zh-tw) Wikipedia pages, gathered between May 1, 2023, and May 7, 2023. Includes both the original HTML format and an auto-converted Markdown version, which has been processed using vinta/pangu.py. 於 2023 年 5 月 1 日至 5 月 7 日間取自維基百科 action=query & prop=extracts API,內容皆與維基百科網站之台灣正體版本一致,沒有繁簡體混雜的問題。 For development… See the full description on the dataset page: https://huggingface.co/datasets/sirzmkk/zh-tw-wikipedia.tabulartext-generation1M<n<10M0 likes267 downloads5mo agoHugging Face24TemryL /tokenized_wikipedia_20220301.en_train_512 Tokenized English Wikipedia Dataset Dataset Description This dataset contains tokenized chunks of text from the English Wikipedia dump of March 1, 2022. Each entry in the dataset represents a chunk of text from Wikipedia, with information about which document and position within the document it comes from. Dataset Creation Source Dataset: Wikipedia (20220301.en) Tokenizer: BERT base uncased Chunk Size: 512 tokens (including special tokens)… See the full description on the dataset page: https://huggingface.co/datasets/TemryL/tokenized_wikipedia_20220301.en_train_512.tabular10M<n<100M0 likes257 downloads2y agoHugging Face25seonglae /wikipedia_token Dataset Card for "wikipedia_token" Token count { '~1024': 5320881, '1024~2048': 693911, '2048~4096': 300935, '4096~8192': 106221, '8192~16384': 30611, '16384~32768': 4812, '32768~65536': 1253, '65536~128000': 46, '128000~': 0 } Text count { '0~1024': 2751539, '1024~2048': 1310778, '2048~4096': 1179150, '4096~8192': 722101, '8192~16384': 329062, '16384~32768': 121237, '32768~65536': 36894, '65536~': 7909 } Token percent { '~1024':… See the full description on the dataset page: https://huggingface.co/datasets/seonglae/wikipedia_token.tabular1M<n<10M0 likes234 downloads3y agoHugging Face26mini97 /filtered_japanese-wikipediatabular1M<n<10M2 likes225 downloads1y agoHugging Face27lewoniewski /wikipedia_quality_wikirankDatasets with WikiRank quality score as of 1 August 2024 for 47 million Wikipedia articles in 55 language versions (also simplified version for each language in separate files). The WikiRank quality score is a metric designed to assess the overall quality of a Wikipedia article. Although its specific algorithm can vary depending on the implementation, the score typically combines several key features of the Wikipedia article. Why It’s Important Enhances Trust: For readers and… See the full description on the dataset page: https://huggingface.co/datasets/lewoniewski/wikipedia_quality_wikirank.tabular10M<n<100M4 likes223 downloads2y agoHugging Face28Reza2kn /Wikipedia-FA-EN-DeepSeek-V4-Flash-0731 Wikipedia Persian to English — DeepSeek V4 Flash 0731 Rolling, machine-generated English translations of Persian Wikipedia articles from Reza2kn/Wikipedia-EN-FA-Accessibility-Bridge, configuration full_articles_fa_without_en. 129,816 translations are currently published in 26 immutable Parquet shards. The target release contains 129,816 translations; five source rows have empty plain_text and are not translated. Shards are published only after 5,000 complete, validated records… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/Wikipedia-FA-EN-DeepSeek-V4-Flash-0731.audiotranslation100K<n<1M1 likes220 downloads24d agoHugging Face29maknee /wikipedia-qwen-4b-clustered-nodiskpq-7to8tabularn<1K0 likes217 downloads6mo agoHugging Face30lianghsun /wikipedia-zh-742M Dataset Card for lianghsun/wikipedia-zh 以繁體中文(zh-tw)語系為主的 Wikipedia 資料集。 Dataset Details Dataset Description 本資料集由自行開發的爬蟲抓取 Wikipedia 上標註繁體中文語系(zh-tw)的文本內容,以確保文本語系是繁體中文。目前 Hugging Face 上標註為繁體中文語系的許多 Wikipedia 資料集,其實並非真正的 Wikipedia 資料集,而是來自 Wikimedia 的內容。本資料集涵蓋範圍廣泛,不僅限於台灣,也包含其他國家的內容,因此可能包含 政治不正確 或 非客觀 的資訊,使用時請謹慎評估。 為便於訓練,同一個 Wikipedia 頁面的語料已被切分為多個子語料,使用者可依需求進行合併處理。好比同為「亞瑟·柯南·道爾」的文本: ... {"text": "同樣在1887年的南海城,他受到了樸茨茅斯文學與哲學學會(Portsmouth Literary and… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/wikipedia-zh-742M.tabulartext-generation1M<n<10M4 likes215 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.