CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kilian-group /phantom-wiki-v0-5-0-predictions Dataset Card for Dataset Name Predictions from https://huggingface.co/datasets/mlcore/phantom-wiki-v050 Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/kilian-group/phantom-wiki-v0-5-0-predictions.tabular100K<n<1M0 likes962 downloads2y agoHugging Face02lflage /wiki-talks Wiki-Talks The Wiki-Talks dataset is a collection of conversational threads extracted from the talk pages on Wikipedia. This dataset captures collaborative dialogue, discussion patterns, and consensus-building among Wikipedia contributors. It is useful for NLP research focused on dialogue, sentiment analysis, and community dynamics. Details Currently due to PyArrow incompatibility to the long recursive structures in the dataset there is an intrinsic incompatibility… See the full description on the dataset page: https://huggingface.co/datasets/lflage/wiki-talks.tabular100K<n<1M1 likes753 downloads2y agoHugging Face03wanng /wikipedia-zh-mnbvc zhwiki-mnbvc 分项目:爬取并处理中文维基百科语料 数据时间:202302-202305 (持续更新) 主项目:MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集 https://github.com/esbatmop/MNBVC 该项目清洗流程主要参考:https://kexue.fm/archives/4176/comment-page-1 并且使用组员开发的去重工具进行数据格式化。 总行数(样本): 10,754,146 一个示例: { "文件名": "cleaned/zhwiki-20230420/folder_0/723712.txt", "是否待查文件": false, "是否重复文件": false, "文件大小": 558, "simhash": 14363740497821204542, "最长段落长度": 142, "段落数": 6, "去重段落数": 6, "低质量段落数": 0, "段落": [ {… See the full description on the dataset page: https://huggingface.co/datasets/wanng/wikipedia-zh-mnbvc.tabulartext-generation1M<n<10M6 likes381 downloads3y agoHugging Face04procesaur /Wiki.korpus Viki korpusi na srpskom i hrvatskom jeziku Sveža verzija, 1. maj 2026! Očišćen i filtriran skup pet projekata: Vikipedija, Vikizvornik, Vikiknjige, Vikivesti i Vikicitati. Preko 670.000 očišćenih članaka, sa preko 310 miliona reči. Svaki dokument je u zasebnoj JSON liniji. Novi metapodaci! Kategorije, broj reči i postotak ćiriličnog teksta Moguće filtiranje skupa po jeziku ili projektu. Wiki corpora in Serbian and… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/Wiki.korpus.tabulartext-generation1M<n<10M1 likes288 downloads2mo agoHugging Face05kierarkia /danbooru-wiki-2026 danbooru-wiki-2026-04-28 About Wiki pages about the danbooru tags on danbooru.donmai.us. The wiki contains the description of each tag. This dataset was collected from the wiki data using danbooru API and then merged with tags data (category_name & post_count). No entries were removed, no matter how many posts they have. The resulting set was manually filtered. Originally some of the lists, tag_group:, about:, help:, howto:, etc, pages were labled as "general"… See the full description on the dataset page: https://huggingface.co/datasets/kierarkia/danbooru-wiki-2026.tabulartext-classification100K<n<1M11 likes283 downloads5mo agoHugging Face06lianghsun /wikipedia-zh-742M Dataset Card for lianghsun/wikipedia-zh 以繁體中文(zh-tw)語系為主的 Wikipedia 資料集。 Dataset Details Dataset Description 本資料集由自行開發的爬蟲抓取 Wikipedia 上標註繁體中文語系(zh-tw)的文本內容,以確保文本語系是繁體中文。目前 Hugging Face 上標註為繁體中文語系的許多 Wikipedia 資料集,其實並非真正的 Wikipedia 資料集,而是來自 Wikimedia 的內容。本資料集涵蓋範圍廣泛,不僅限於台灣,也包含其他國家的內容,因此可能包含 政治不正確 或 非客觀 的資訊,使用時請謹慎評估。 為便於訓練,同一個 Wikipedia 頁面的語料已被切分為多個子語料,使用者可依需求進行合併處理。好比同為「亞瑟·柯南·道爾」的文本: ... {"text": "同樣在1887年的南海城,他受到了樸茨茅斯文學與哲學學會(Portsmouth Literary and… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/wikipedia-zh-742M.tabulartext-generation1M<n<10M4 likes218 downloads2y agoHugging Face07maknee /wikipedia-qwen-4b-clustered-nodiskpq-7to8tabularn<1K0 likes217 downloads6mo agoHugging Face08iperbole /wiki-to-rcqa-italian Wiki-to-RCQA - Italian (IT) tabulartext-generation1M<n<10M0 likes186 downloads18d agoHugging Face09superAVTR /wikipedia_en_512_for_pretraining Cleaned Wikipedia 512 Pretraining Dataset Dataset Description This dataset is a cleaned version of the Hugging Face dataset lucadiliello/wikipedia_512_pretraining. The original dataset contains English Wikipedia text prepared for language-model pretraining. This derivative version applies additional filtering intended to remove malformed, duplicated, or incomplete training samples while otherwise preserving the original text. Source Original… See the full description on the dataset page: https://huggingface.co/datasets/superAVTR/wikipedia_en_512_for_pretraining.tabular1M<n<10M0 likes143 downloads7d agoHugging Face10hallisky /wikiMIA-2024-hard WikiMIA-2024 Hard Dataset Dataset Description WikiMIA_2024 Hard is a challenging dataset for membership inference attacks intorduced in the paper "The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage" containing temporal Wikipedia articles with different versions based on date cutoffs. This dataset is designed to evaluate the robustness of privacy-preserving machine learning models against sophisticated membership inference techniques. It… See the full description on the dataset page: https://huggingface.co/datasets/hallisky/wikiMIA-2024-hard.tabulartext-classification1K<n<10K0 likes137 downloads1y agoHugging Face11procesaur /Wiki.sl Wiki korpusi v slovenščini Sveža različica, 1. maj 2026! Očiščen in filtriran nabor štirih projektov: Wikipedia, Wikivir, Wikiknjige in Wikinavedki. Več kot 146.000 kuriranih člankov z več kot 130 milijoni besed. Vsak dokument je v ločeni vrstici JSON. Novi metapodatki! Kategorije, število besed (in odstotek ciriličnega besedila) Možnost filtriranja nabora po jeziku ali projektu. Wiki corpora in Slovenian language… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/Wiki.sl.tabulartext-generation100K<n<1M0 likes134 downloads2mo agoHugging Face12pere /wiki_paragraphs_norwegian WIKI Paragraphs Norwegian A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format. Features Multiple splits for different use cases Random shuffle with Fisher-Yates algorithm Structured format with text and metadata Size-varied validation/test sets (100 to 10k samples) Splits Overview Split Name Samples Typical Usage train 1,000,000 Primary training data validation 10,000 Standard… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_norwegian.tabulartext-generation1M<n<10M0 likes131 downloads2y agoHugging Face13arnomatic /german-wikipedia-clean-2tabular1M<n<10M1 likes125 downloads11mo agoHugging Face14kilian-group /phantom-wiki-v0.3-null Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Format based on this dataset: https://huggingface.co/rag-datasets Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/kilian-group/phantom-wiki-v0.3-null.tabular10K<n<100K0 likes123 downloads2y agoHugging Face15procesaur /Wiki.mk Вики корпуси на македонски јазик Свежа верзија, 1 мај 2026! Исчистен и филтриран сет од три проекти: Википедија, Викиизвор и Викикниги Над 100.000 курирани статии, со над 50 милиони зборови. Секој документ е во посебна JSON линија. Нови метаподатоци! Категории, број на зборови и процент на кириличен текст Можно е да се филтрира сет по јазик или проект. Wiki corpora in Macedonian Fresh version, 1. May 2026!… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/Wiki.mk.tabulartext-generation100K<n<1M0 likes115 downloads2mo agoHugging Face16CJJones /Wikipedia_RAG_QA_Classification 🏛️ Wikipedia RAG QA Dataset for Retrieval-Augmented Generation Training 📊 Dataset Description This dataset contains 300,000+ validated model-generated responses to Wikipedia content, specifically designed for Retrieval-Augmented Generation (RAG) applications and SQL database insertion tasks. Generated by Jeeney AI Reloaded 207M GPT with specialized RAG tuning. 🖥️ Demo Interface: Discord Live Chat Demo on Discord: https://discord.gg/Xe9tHFCS9h The full CJ… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Wikipedia_RAG_QA_Classification.tabulartext-generation100K<n<1M1 likes112 downloads6mo agoHugging Face17pere /wiki_paragraphs_english WIKI Paragraphs English A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format. Features Multiple splits for different use cases Random shuffle with Fisher-Yates algorithm Structured format with text and metadata Size-varied validation/test sets (100 to 10k samples) Splits Overview Split Name Samples Typical Usage train 1,000,000 Primary training data validation 10,000 Standard validation… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_english.tabulartext-generation1M<n<10M0 likes84 downloads2y agoHugging Face18malaiwah /qfs-smollm2-135m-wikitext2-native-v1 HF workflow d3dc69602aeb981f06bd9f4c726937f9 A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-native-bf16. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-native-v1.tabularn<1K0 likes67 downloads16d agoHugging Face19pere /nor_wiki_reasoning_english_v4tabular1K<n<10K0 likes63 downloads2y agoHugging Face20procesaur /Wiki.bg Уики корпуси на македонски и български език Нова версия, 1 май 2026 г! Почистен и филтриран набор от пет проекта: Уикипедия, Уикиизточник, Уикикниги, Уикиновини и Уикицитат. Над 240 000 курирани статии, с над 100 милиона думи. Всеки документ е на отделен JSON ред. Нови метаданни! Категории, брой думи и процент на кирилица Възможно е да филтрирате набора по език или проект. Wiki corpora in Bulgarian Fresh… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/Wiki.bg.tabulartext-generation100K<n<1M0 likes63 downloads2mo agoHugging Face21festr2 /glm51-kld-reference-logits-wikitext-ctx2048-s512-20260517 --- license: other pretty_name: GLM-5.1 KLD Reference Logits WikiText ctx2048 s512 tags: - logits - kld - glm-5.1 - vllm - b12x --- # GLM-5.1 KLD Reference Logits Public cache of the reference logits used for GLM-5.1 NVFP4 / mixed FP8_PB_WO KLD evaluation. These files are generated logits, not model weights. They are stored as `logits_*.safetensors` with one tensor named `logits`, shape `(2047, 154880)`, dtype `float32`. ##… See the full description on the dataset page: https://huggingface.co/datasets/festr2/glm51-kld-reference-logits-wikitext-ctx2048-s512-20260517.tabularn<1K0 likes60 downloads4mo agoHugging Face22artur-muratov /wikiSQL-kk-datasetimage10K<n<100K0 likes57 downloads2y agoHugging Face23trentmkelly /soyjak-wiki Soyjak Wiki A full dump of Soyjak Wiki, a MediaWiki-based encyclopedia documenting soyjak memes, variants, culture, communities, and related internet history. The dump includes all 8,039 pages (2,927 articles and 5,112 redirects) with raw wikitext markup preserved. Columns Column Type Description title string Page title page_id int MediaWiki page ID revision_id int Revision ID of the exported version timestamp string Last edit timestamp (ISO 8601)… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/soyjak-wiki.tabular1K<n<10K1 likes56 downloads7mo agoHugging Face24proxectonos /wikipedia_multiple_choice_qa Galician and Portuguese Multiple-Choice QA Instruction Subsets Dataset description This dataset contains two instruction-tuning subsets for multiple-choice question answering in Galician and Portuguese: gl_wikipedia_multiple_choice_qa (1,486 instances) pt_wikipedia_multiple_choice_qa (547 instances) Both subsets are reformatted versions of QA data originally included in the cpt_instruction_datasets collection, adapted here as standalone instruction-style datasets. Each… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/wikipedia_multiple_choice_qa.tabulartext-generation1K<n<10K1 likes55 downloads5mo agoHugging Face25malaiwah /qfs-smollm2-135m-wikitext2-gptq-g32-v1 HF workflow 32c6ab05b0ceab1cecdceda838846388 A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-gptq-int4-g32. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-gptq-g32-v1.tabularn<1K0 likes55 downloads16d agoHugging Face26omron-sinicx /wiki2023_plus Overview The dataset includes the description from Wikipedia and categories of films published in 2023. This dataset is used to evaluate the ability of LLM to memorize and extract information described in the document. See "Where is the Answer? An Empirical Study of Positional Bias for Parametric Knowledge Extraction in Language Model (NAACL2025 Long paper)" for how we use this dataset for training and evaluation. Data Split film_doc_all.jsonl includes lines of… See the full description on the dataset page: https://huggingface.co/datasets/omron-sinicx/wiki2023_plus.tabular1K<n<10K1 likes52 downloads1y agoHugging Face27artur-muratov /wikiSQL-ru-datasetimage10K<n<100K0 likes51 downloads2y agoHugging Face28kaan39 /turkish-wikipedia-dataset-clean Turkish Wikipedia Dataset A cleaned and structured Turkish Wikipedia dataset designed for Turkish language model pretraining, continued pretraining, research, and NLP experiments. The dataset consists of articles collected from the Turkish Wikipedia (tr.wikipedia.org) and processed into a machine-readable format while preserving important source metadata. Dataset Summary Language: Turkish (tr) Source: Turkish Wikipedia Domain: General knowledge / encyclopedia… See the full description on the dataset page: https://huggingface.co/datasets/kaan39/turkish-wikipedia-dataset-clean.tabular100K<n<1M0 likes51 downloads1mo agoHugging Face29Ram-G /Wiki_Faiss_Indexes dataset_info: features: - name: text dtype: string - name: embeddings dtype: float32 shape: [384] configs: - config_name: default data_files: "*.parquet" Wikipedia IVF-OPQ-PQ Vector Database (GPU-Optimized) A high-performance, GPU-accelerated FAISS vector database built from Wikipedia articles with pre-computed embeddings. This dataset contains approximately 35 million Wikipedia articles with 384-dimensional embeddings using the all-MiniLM-L6-v2… See the full description on the dataset page: https://huggingface.co/datasets/Ram-G/Wiki_Faiss_Indexes.tabularfeature-extractionn<1K1 likes50 downloads1y agoHugging Face30malaiwah /qfs-smollm2-135m-wikitext2-gptq-g64-v1 HF workflow 73f0a12a901c7368794a3a886f55b675 A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-gptq-int4-g64. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-gptq-g64-v1.tabularn<1K0 likes49 downloads16d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.