CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01common-pile /wikimedia_filtered Wikimedia Description Official Wikimedia wikis are released under a CC BY-SA license. We downloaded the official database dumps from March 2025 of the English-language wikis that are directly managed by the Wikimedia Foundation. These database dumps include the wikitext—MediaWiki’s custom markup language—for each page as well as talk pages, where editors discuss changes made to a page. We only use the most recent version of each page. We converted wikitext to plain text… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/wikimedia_filtered.texttext-generation10M<n<100M8 likes2.8k downloads1y agoHugging Face020xDing /wikipedia-cn-20230720-filtered本数据集基于中文维基2023年7月20日的dump存档。作为一项以数据为中心的工作,本数据集仅保留了 254,547条 质量较高的词条内容。具体而言: 过滤了Template, Category, Wikipedia, File, Topic, Portal, MediaWiki, Draft, Help等特殊类型的词条 使用启发式的方法和自有的NLU模型过滤了一部分质量较低的词条 过滤了一部分内容较为敏感或存在争议性的词条。 进行了简繁转换和习惯用词转换,确保符合中国大陆地区的习惯用词。 This dataset is based on the Chinese Wikipedia dump archive from July 20th, 2023. As a data-centric effort, the dataset retains 254,574 high-quality entries. Specifically: Entries of special types such as Template, Category, Wikipedia, File, Topic… See the full description on the dataset page: https://huggingface.co/datasets/0xDing/wikipedia-cn-20230720-filtered.texttext-generation100K<n<1M173 likes2.2k downloads3y agoHugging Face03common-pile /wikiteam Wikiteam Description There are many wikis on the internet that are not managed by the Wikimedia foundation, but do use their MediaWiki software to power their wiki. Many of these wikis have been archived by wikiteam, a collection of volunteers that create unofficial database dumps of wikis and upload them to the Internet Archive. We download all dumps made by wikiteam when the metadata indicates the wiki was licensed under CC BY, CC BY-SA, or released into the public… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/wikiteam.texttext-generation100M<n<1B3 likes1.5k downloads1y agoHugging Face04common-pile /wikimedia Wikimedia Description Official Wikimedia wikis are released under a CC BY-SA license. We downloaded the official database dumps from March 2025 of the English-language wikis that are directly managed by the Wikimedia foundation. These database dumps include the wikitext—Mediawiki’s custom markup language—for each page as well as talk pages, where editors discuss changes made for a page. We only use the most recent version of each page. We converted wikitext to plain… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/wikimedia.texttext-generation10M<n<100M6 likes932 downloads1y agoHugging Face05Podtech /llm-jp-corpus-v4-ja_wiki llm-jp-corpus-v4 — ja_wiki Mirror of the ja/ja_wiki sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_wiki Files: 6 × jsonl.gz (1.9 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository. License CC BY-SA 3.0 — inherited from the… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_wiki.texttext-generation1M<n<10M0 likes879 downloads2mo agoHugging Face06agentlans /wikipedia-paragraphs Wikipedia Paragraph Samples Dataset Description This dataset contains paragraphs extracted from randomly selected English Wikipedia articles. It provides a diverse sample of Wikipedia content across various topics. Dataset Details Name: Wikipedia Paragraph Samples Version: 1.0 Date Created: 2024-08-20 Language: English Format: JSONLines Contents Each line in the dataset represents a single paragraph and contains two fields: Title of the Wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraphs.texttext-classification10K<n<100K3 likes436 downloads2y agoHugging Face07claran /m2d2-wiki-decon Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/claran/m2d2-wiki-decon.texttext-generation1M<n<10M0 likes405 downloads2y agoHugging Face08wanng /wikipedia-zh-mnbvc zhwiki-mnbvc 分项目:爬取并处理中文维基百科语料 数据时间:202302-202305 (持续更新) 主项目:MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集 https://github.com/esbatmop/MNBVC 该项目清洗流程主要参考:https://kexue.fm/archives/4176/comment-page-1 并且使用组员开发的去重工具进行数据格式化。 总行数(样本): 10,754,146 一个示例: { "文件名": "cleaned/zhwiki-20230420/folder_0/723712.txt", "是否待查文件": false, "是否重复文件": false, "文件大小": 558, "simhash": 14363740497821204542, "最长段落长度": 142, "段落数": 6, "去重段落数": 6, "低质量段落数": 0, "段落": [ {… See the full description on the dataset page: https://huggingface.co/datasets/wanng/wikipedia-zh-mnbvc.tabulartext-generation1M<n<10M6 likes390 downloads3y agoHugging Face09common-pile /wikiteam_filtered Wikiteam Description There are many wikis on the internet that are not managed by the Wikimedia Foundation, but do use their MediaWiki software to power their wiki. Many of these wikis have been archived by Wikiteam, a collection of volunteers that create unofficial database dumps of wikis and upload them to the Internet Archive. We download all dumps made by Wikiteam when the metadata indicates the wiki was licensed under CC BY, CC BY-SA, or released into the public… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/wikiteam_filtered.texttext-generation10M<n<100M2 likes327 downloads1y agoHugging Face10marin-community /wikipedia-markdown Marin Markdownified Wikipedia Markdownified Wikipedia is a large-scale, pre-processed version of the English Wikipedia Enterprise HTML dump consisting of 8.59B tokens. The corpus has been converted to clean, section-aware Markdown for language-model training. Value Tokens 8 587 224 558 Primary source https://dumps.wikimedia.org/other/enterprise_html/runs/20241201/enwiki-NS0-20241201-ENTERPRISE-HTML.json.tar.gz File format JSONL License CC-BY-SA 4.0 (mirrors… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/wikipedia-markdown.texttext-generation1M<n<10M7 likes310 downloads1y agoHugging Face11procesaur /Wiki.korpus Viki korpusi na srpskom i hrvatskom jeziku Sveža verzija, 1. maj 2026! Očišćen i filtriran skup pet projekata: Vikipedija, Vikizvornik, Vikiknjige, Vikivesti i Vikicitati. Preko 670.000 očišćenih članaka, sa preko 310 miliona reči. Svaki dokument je u zasebnoj JSON liniji. Novi metapodaci! Kategorije, broj reči i postotak ćiriličnog teksta Moguće filtiranje skupa po jeziku ili projektu. Wiki corpora in Serbian and… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/Wiki.korpus.tabulartext-generation1M<n<10M1 likes281 downloads2mo agoHugging Face12kierarkia /danbooru-wiki-2026 danbooru-wiki-2026-04-28 About Wiki pages about the danbooru tags on danbooru.donmai.us. The wiki contains the description of each tag. This dataset was collected from the wiki data using danbooru API and then merged with tags data (category_name & post_count). No entries were removed, no matter how many posts they have. The resulting set was manually filtered. Originally some of the lists, tag_group:, about:, help:, howto:, etc, pages were labled as "general"… See the full description on the dataset page: https://huggingface.co/datasets/kierarkia/danbooru-wiki-2026.tabulartext-classification100K<n<1M11 likes269 downloads5mo agoHugging Face13lukasthede /WikiBigEdit 📚 WikiBigEdit Paper (EasyEdit2 Framework): EasyEdit2: An Easy-to-use Steering Framework for Editing Large Language Models Code (EasyEdit Framework): https://github.com/zjunlp/EasyEdit Project Page (EasyEdit2 Framework): https://zjunlp.github.io/project/EasyEdit2/ Paper (WikiBigEdit Dataset): Understanding the Limits of Lifelong Knowledge Editing in LLMs 🌟 Overview WikiBigEdit is a large-scale benchmark designed to assess the ability of language models to integrate… See the full description on the dataset page: https://huggingface.co/datasets/lukasthede/WikiBigEdit.textquestion-answering100K<n<1M4 likes252 downloads1y agoHugging Face14lianghsun /wikipedia-zh-742M Dataset Card for lianghsun/wikipedia-zh 以繁體中文(zh-tw)語系為主的 Wikipedia 資料集。 Dataset Details Dataset Description 本資料集由自行開發的爬蟲抓取 Wikipedia 上標註繁體中文語系(zh-tw)的文本內容,以確保文本語系是繁體中文。目前 Hugging Face 上標註為繁體中文語系的許多 Wikipedia 資料集,其實並非真正的 Wikipedia 資料集,而是來自 Wikimedia 的內容。本資料集涵蓋範圍廣泛,不僅限於台灣,也包含其他國家的內容,因此可能包含 政治不正確 或 非客觀 的資訊,使用時請謹慎評估。 為便於訓練,同一個 Wikipedia 頁面的語料已被切分為多個子語料,使用者可依需求進行合併處理。好比同為「亞瑟·柯南·道爾」的文本: ... {"text": "同樣在1887年的南海城,他受到了樸茨茅斯文學與哲學學會(Portsmouth Literary and… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/wikipedia-zh-742M.tabulartext-generation1M<n<10M4 likes222 downloads2y agoHugging Face15Blaze7451 /Wiki-zhtw-20250601 Dataset Card for Wiki-zhtw-20250601 Dataset Description This dataset is derived from the Chinese‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim, converted to Markdown format via regular‑expression post‑processing, and finally converted from Simplified to Traditional Chinese using OpenCC. texttext-generation1M<n<10M1 likes198 downloads1y agoHugging Face16Mungus451 /verified_wiki_historian_the_beatles_anthology_dataset_active Verified-Wiki-Historian: The Beatles Anthology Verified-Wiki-Historian (The Beatles Anthology) is a refined, citation-grounded instruction dataset for Beatles-specific historical question answering, summarization, and supervised fine-tuning. This dataset is a cleaned and rebuilt refinement of: Mungus451/verified_wiki_historian_the_beatles_anthology_dataset_active The current release contains 4,000 instruction records focused on Beatles history, recording sessions, release… See the full description on the dataset page: https://huggingface.co/datasets/Mungus451/verified_wiki_historian_the_beatles_anthology_dataset_active.textquestion-answering1K<n<10K1 likes197 downloads3mo agoHugging Face17its5Q /wikireading Dataset Card for Wikireading This is a dataset of book chapters scraped from a Russian website called Wikireading. Dataset Details Dataset Description Wikireading is a collection of non-fiction educational books in various domains: Biology, Art, History, Religion and much more. The books are highly educational and provide vast knowledge in different domains, making this dataset a good choice for pretraining. The resulting dataset contains ~26M rows, which in… See the full description on the dataset page: https://huggingface.co/datasets/its5Q/wikireading.texttext-generation1M<n<10M9 likes195 downloads2y agoHugging Face18agentlans /wikipedia-first-paragraphtexttext-classification10M<n<100M0 likes189 downloads1y agoHugging Face19iperbole /wiki-to-rcqa-italian Wiki-to-RCQA - Italian (IT) tabulartext-generation1M<n<10M0 likes187 downloads18d agoHugging Face20costadev00 /openai-terra-batch-wiki-brazil-1000-partial-20260724-01 OpenAI Terra Batch — Wikipédia PT-BR (run parcial) Checkpoint publicável de uma execução real e interrompida do fluxo document_task_matrix. A execução planejou gerar uma matriz de 1.000 documentos da Wikipédia em português por 25 tasks canônicas usando a Responses API Batch e o modelo gpt-5.6-terra. Este repositório não representa a conclusão dos 25.000 pares planejados. Ele contém somente os 1.282 candidatos aceitos após a reconciliação offline de todos os resultados Batch já… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/openai-terra-batch-wiki-brazil-1000-partial-20260724-01.texttext-generation1K<n<10K0 likes171 downloads2mo agoHugging Face21gouwsxander /wikipedia-human-ai Wikipedia Human/AI A selection of ~10,000 paragraphs from Wikipedia, along with rewritten text by GPT 5 Nano. texttext-classification1K<n<10K10 likes159 downloads10mo agoHugging Face22pere /wiki_paragraphs_norwegian WIKI Paragraphs Norwegian A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format. Features Multiple splits for different use cases Random shuffle with Fisher-Yates algorithm Structured format with text and metadata Size-varied validation/test sets (100 to 10k samples) Splits Overview Split Name Samples Typical Usage train 1,000,000 Primary training data validation 10,000 Standard… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_norwegian.tabulartext-generation1M<n<10M0 likes138 downloads2y agoHugging Face23procesaur /Wiki.sl Wiki korpusi v slovenščini Sveža različica, 1. maj 2026! Očiščen in filtriran nabor štirih projektov: Wikipedia, Wikivir, Wikiknjige in Wikinavedki. Več kot 146.000 kuriranih člankov z več kot 130 milijoni besed. Vsak dokument je v ločeni vrstici JSON. Novi metapodatki! Kategorije, število besed (in odstotek ciriličnega besedila) Možnost filtriranja nabora po jeziku ali projektu. Wiki corpora in Slovenian language… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/Wiki.sl.tabulartext-generation100K<n<1M0 likes130 downloads2mo agoHugging Face24Blaze7451 /Wiki-ja-20250601 Dataset Card for Wiki-ja-20250601 Dataset Description This dataset is derived from the Japan‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim and converted to Markdown format via regular‑expression post‑processing. texttext-generation100K<n<1M0 likes126 downloads1y agoHugging Face25ikedachin /imabari_wiki_qa_v4_validated_w_reasoning_effort_qwen38 Imabari Wiki QA v4 Validated with Reasoning Effort — Qwen3.8 概要 / Overview 日本語・今治弁のQAを用いて、reasoning effort に応じた思考文の生成を学習するための教師ありファインチューニング(SFT)用データセットです。Imabari Wiki QA v4 Validated の質問と回答を保持し、元記事の文脈を参照して思考文を再生成しています。 This dataset supports supervised fine-tuning (SFT) of reasoning-effort-conditioned explanations using Japanese QA with Imabari dialect expressions. Questions and answers from Imabari Wiki QA v4 Validated are preserved, while reasoning text is… See the full description on the dataset page: https://huggingface.co/datasets/ikedachin/imabari_wiki_qa_v4_validated_w_reasoning_effort_qwen38.textquestion-answering10K<n<100K0 likes122 downloads7d agoHugging Face26ikedachin /imabari_wiki_qa_v4_validated Imabari QA v4 — Validated Dataset Summary Imabari QA v4 — Validated is a Japanese question-answering dataset for supervised fine-tuning (SFT). This dataset combines two independently curated variants of the Imabari QA v4 synthetic reasoning dataset: ikedachin/imabari_qa_v4_program_validated ikedachin/imabari_qa_v4_human_validated Both datasets are derived from: Source corpus: ikedachin/imabari_wiki_cpt_v3 QA generation model: Qwen3.8-27B-NVFP4 The combined… See the full description on the dataset page: https://huggingface.co/datasets/ikedachin/imabari_wiki_qa_v4_validated.textquestion-answering1K<n<10K0 likes116 downloads14d agoHugging Face27CJJones /Wikipedia_RAG_QA_Classification 🏛️ Wikipedia RAG QA Dataset for Retrieval-Augmented Generation Training 📊 Dataset Description This dataset contains 300,000+ validated model-generated responses to Wikipedia content, specifically designed for Retrieval-Augmented Generation (RAG) applications and SQL database insertion tasks. Generated by Jeeney AI Reloaded 207M GPT with specialized RAG tuning. 🖥️ Demo Interface: Discord Live Chat Demo on Discord: https://discord.gg/Xe9tHFCS9h The full CJ… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Wikipedia_RAG_QA_Classification.tabulartext-generation100K<n<1M1 likes113 downloads6mo agoHugging Face28procesaur /Wiki.mk Вики корпуси на македонски јазик Свежа верзија, 1 мај 2026! Исчистен и филтриран сет од три проекти: Википедија, Викиизвор и Викикниги Над 100.000 курирани статии, со над 50 милиони зборови. Секој документ е во посебна JSON линија. Нови метаподатоци! Категории, број на зборови и процент на кириличен текст Можно е да се филтрира сет по јазик или проект. Wiki corpora in Macedonian Fresh version, 1. May 2026!… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/Wiki.mk.tabulartext-generation100K<n<1M0 likes112 downloads2mo agoHugging Face29Blaze7451 /Wiki-ko-20250601 Dataset Card for Wiki-ko-20250601 Dataset Description This dataset is derived from the Korea‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim and converted to Markdown format via regular‑expression post‑processing. texttext-generation1M<n<10M0 likes99 downloads1y agoHugging Face30Blaze7451 /Wiki-vi-20250601 Dataset Card for Wiki-vi-20250601 Dataset Description This dataset is derived from the Vietnam‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim and converted to Markdown format via regular‑expression post‑processing. texttext-generation1M<n<10M1 likes97 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.