CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Intelligent-Internet /wikipedia_en wikipedia_en This is a curated Wikipedia English dataset for use with the II-Commons project. Dataset Details Dataset Description This dataset comprises a curated Wikipedia English pages. Data sourced directly from the official English Wikipedia database dump. We extract the pages, chunk them into smaller pieces, and embed them using Snowflake/snowflake-arctic-embed-m-v2.0. All vector embeddings are 16-bit half-precision vectors optimized for cosine indexing… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/wikipedia_en.tabularfeature-extraction10M<n<100M2 likes1.2k downloads1y agoHugging Face02EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1007 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007.tabular10K<n<100K0 likes890 downloads25d agoHugging Face03EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1004 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1004 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1004.tabular10K<n<100K0 likes854 downloads25d agoHugging Face04EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1006 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006.tabular10K<n<100K0 likes846 downloads25d agoHugging Face05vtasca /wikipedia-pageviews Wikipedia Article Pageviews This repository automatically fetches and aggregates the 100 most popular Wikipedia articles by pageviews - creating a dataset that enables tracking trending topics on Wikipedia. It works by polling the WikiMedia API on a daily basis and fetching the top 100 most popular articles from two days ago. The fetcher runs in a scheduled GitHub Actions workflow, which is available here. The dataset begins in the year 2016 and the textual data is presented as… See the full description on the dataset page: https://huggingface.co/datasets/vtasca/wikipedia-pageviews.tabularfeature-extraction100K<n<1M1 likes842 downloads38m agoHugging Face06EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1008 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1008 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1008.tabular10K<n<100K0 likes839 downloads25d agoHugging Face07Metin /WikiRAG-TR Dataset Summary WikiRAG-TR is a dataset of 6K (5999) question and answer pairs which synthetically created from introduction part of Turkish Wikipedia Articles. The dataset is created to be used for Turkish Retrieval-Augmented Generation (RAG) tasks. Dataset Information Number of Instances: 5999 (5725 synthetically generated question-answer pairs, 274 augmented negative samples) Dataset Size: 20.5 MB Language: Turkish Dataset License: apache-2.0 Dataset Category:… See the full description on the dataset page: https://huggingface.co/datasets/Metin/WikiRAG-TR.tabularquestion-answering1K<n<10K36 likes491 downloads2y agoHugging Face08google /WikiProfile WikiProfile WikiProfile is a factual knowledge benchmark for evaluating how well language models encode and recall factual knowledge. It comprises 2,150 facts, each paired with 10 questions, for a total of 21,500 question instances. Each fact is grounded in the first paragraph (summary) of an English Wikipedia page and is defined as a proposition between two entities, a subject and an object (e.g., "Oasis played their first gig at the Boardwalk club" → subject: Oasis, object:… See the full description on the dataset page: https://huggingface.co/datasets/google/WikiProfile.tabularquestion-answering1K<n<10K20 likes467 downloads3mo agoHugging Face09wikilee /ADFA_Mappingtabular10M<n<100M0 likes460 downloads5y agoHugging Face10EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1005 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1005 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1005.tabular10K<n<100K0 likes422 downloads25d agoHugging Face11mgmacleod /wikidata1 Wikipedia N-Link Basins A novel graph-theoretic analysis of Wikipedia's internal link structure, revealing deterministic "basins of attraction" under N-link traversal rules. Dataset Description This dataset demonstrates that Wikipedia's 17.9 million pages partition into coherent regions when following a simple rule: from any page, always follow the Nth link. Every page eventually reaches a cycle, and pages sharing the same terminal cycle form a basin of attraction.… See the full description on the dataset page: https://huggingface.co/datasets/mgmacleod/wikidata1.tabulargraph-mln<1K0 likes263 downloads9mo agoHugging Face12lewoniewski /wikipedia_quality_wikirankDatasets with WikiRank quality score as of 1 August 2024 for 47 million Wikipedia articles in 55 language versions (also simplified version for each language in separate files). The WikiRank quality score is a metric designed to assess the overall quality of a Wikipedia article. Although its specific algorithm can vary depending on the implementation, the score typically combines several key features of the Wikipedia article. Why It’s Important Enhances Trust: For readers and… See the full description on the dataset page: https://huggingface.co/datasets/lewoniewski/wikipedia_quality_wikirank.tabular10M<n<100M4 likes225 downloads2y agoHugging Face13MongoDB /cosmopedia-wikihow-chunked Overview This dataset is a chunked version of a subset of data in the Cosmopedia dataset curated by Hugging Face. Specifically, we have only used a subset of Wikihow articles from the Cosmopedia dataset, and each article has been split into chunks containing no more than 2 paragraphs. Dataset Structure Each record in the dataset represents a chunk of a larger article, and contains the following fields: doc_id: A unique identifier for the parent article chunk_id: A unique… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/cosmopedia-wikihow-chunked.tabularquestion-answering1M<n<10M9 likes207 downloads3y agoHugging Face14anhaidgroup /polaris-wikitables-v2 WikiTables v2 WikiTables is one of six datasets in Polaris: Learning to Generate Table Descriptions from Retrieval Feedback, alongside aw, arctic, lter, ecir, and wtr. It holds 3,361 tables scraped from Wikipedia articles and 57 keyword queries over them. For each query–table pair, a person scored how well that table answers that query; those scores are the relevance judgments, and they live in qrels.csv. The tables have no names. What describes a table is its column names and… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-wikitables-v2.tabulartext-retrieval1K<n<10K0 likes202 downloads1mo agoHugging Face15aadityaubhat /GPT-wiki-intro GPT Wiki Intro Overview Dataset for training models to classify human written vs GPT/ChatGPT generated text. This dataset contains Wikipedia introductions and GPT (Curie) generated introductions for 150k topics. Prompt used for generating text 200 word wikipedia style introduction on '{title}' {starter_text} where title is the title for the wikipedia page, and starter_text is the first seven words of the wikipedia introduction. Here's an example of prompt used to… See the full description on the dataset page: https://huggingface.co/datasets/aadityaubhat/GPT-wiki-intro.tabulartext-classification100K<n<1M27 likes198 downloads3y agoHugging Face16yashassnadig /wikimovies Wikipedia Movies Dataset Dataset Description This dataset contains 58.1k movie information scraped from Wikipedia's "List of films" pages. The data includes basic movie metadata, infobox information, and introductory text from individual movie Wikipedia pages. NOTE: This is an uncleaned dataset containing raw scraped data. The content is sourced from Wikipedia and is not owned by the dataset creator. All content remains under Wikipedia's licensing terms.… See the full description on the dataset page: https://huggingface.co/datasets/yashassnadig/wikimovies.tabulartext-classification10K<n<100K1 likes176 downloads1y agoHugging Face17EleutherAI /bergson-wikitext-gpt2-leaderboard-bank bergson leaderboard: retrain banks, scores and LDS/QLD results (WikiText GPT-2) Everything behind the numbers on the bergson leaderboard, for the model at EleutherAI/bergson-wikitext-gpt2-leaderboard. path what it is bank/ the LDS ground truth: 100 random leave-1%-out subsets of the 4,608 training chunks (subsets.json) and each subset's measured loss change on the 50 test queries (validation.csv) random/retrained/{base,subset_0..99} the retrained models themselves… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/bergson-wikitext-gpt2-leaderboard-bank.tabular10K<n<100K0 likes168 downloads8d agoHugging Face18aisuko /simple_english_wikipediaOnly for the reaseaching usage. The original data from http://sbert.net/datasets/simplewiki-2020-11-01.jsonl.gz. We use nq_distilbert-base-v1 model encode all the data to the PyTorch Tensors. And normalize the embeddings by using sentence_transformers.util.normalize_embeddings. How to use See notebook Wikipedia Q&A Retrieval-Semantic Search Installing the package !pip install sentence-transformers==2.3.1 The converting process # the whole process takes… See the full description on the dataset page: https://huggingface.co/datasets/aisuko/simple_english_wikipedia.tabular100K<n<1M0 likes100 downloads3y agoHugging Face19lgelauff /wikipedia-image-requests Daily image-request counts for Wikipedia article images Daily counts of how often each of ~1.19 million Wikipedia article images was requested from Wikimedia's image servers, attributed to the wiki whose page the request came from, together with the article each image appears on. 406,263,728 daily observations covering 1,186,150 image files across 534,067 articles in nine Wikipedias, from 2025-09-01 to 2026-09-05 (370 days). Every file in the set appears on exactly one… See the full description on the dataset page: https://huggingface.co/datasets/lgelauff/wikipedia-image-requests.tabular100M<n<1B0 likes89 downloads8d agoHugging Face20muset-ai /Wiki_Live_Challenge Wiki Live Challenge Dataset [English | 中文] English 📖 Dataset Overview This is the official dataset accompanying the Wiki Live Challenge benchmark. It contains Wikipedia Good Articles (GAs) as ground truth and research articles generated by leading deep research AI systems. Wiki Live Challenge is the first live benchmark for evaluating Deep Research Agents (DRAs) on their ability to generate Wikipedia-quality articles. Unlike static benchmarks, Wiki Live… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/Wiki_Live_Challenge.tabulartext-generationn<1K1 likes72 downloads8mo agoHugging Face21anhaidgroup /polaris-wikitables-v1 WikiTables v1 WikiTables is one of six datasets in Polaris: Learning to Generate Table Descriptions from Retrieval Feedback, alongside aw, arctic, lter, ecir, and wtr. It holds 3,361 tables scraped from Wikipedia articles and 57 keyword queries over them. For each query–table pair, a person scored how well that table answers that query; those scores are the relevance judgments, and they live in qrels.csv. The tables have no names. What describes a table is its column names and… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-wikitables-v1.tabulartext-retrieval1K<n<10K0 likes60 downloads1mo agoHugging Face22lewoniewski /wikipedia-citation-indexDataset with citation indexes as of 1 August 2024 for 47 million Wikipedia articles in 55 language versions. Research: ArXiv tabular10M<n<100M0 likes48 downloads1y agoHugging Face23ashmib /wikivoyage-eu-city-embeddings Dataset Card for Dataset Name This dataset comprises abstracts from Wikivoyage for 160 European cities along with their corresponding country names, coordinates, and populations. The embeddings are derived from the GTE-Large model, incorporating data from the city, country, population, and abstract columns. Dataset Sources Wikivoyage data World cities database tabularn<1K0 likes37 downloads3y agoHugging Face24zyzw10086 /moonlighter-2-wiki-data Moonlighter 2 relic prices and weapons dataset An open, versioned export from www.moonlighter2.wiki, an independent and unofficial Moonlighter 2 wiki published as The Endless Ledger. This release contains two small, research-friendly tables: Relic prices: 159 published relics, including 154 rows with sourced prices and 5 rows whose unknown prices are intentionally left blank. Weapons: 30 published weapons with weapon type, effects, upgrade information, verification state and… See the full description on the dataset page: https://huggingface.co/datasets/zyzw10086/moonlighter-2-wiki-data.tabularn<1K0 likes37 downloads9d agoHugging Face25kornwtp /thai-wikiqa-tha-qaretrievalref: https://aiforthai.in.th/ tabular10K<n<100K0 likes36 downloads2y agoHugging Face26julia-lukasiewicz-pater /small-GPT-wiki-intro-features Small-GPT-wiki-intro-features dataset This dataset is based on aadityaubhat/GPT-wiki-intro. It contains 100k randomly selected texts (50k from Wikipedia and 50k generated by ChatGPT). For each text, various complexity measures were calculated, including e.g. readibility, lexical richness etc. It can be used for text classification or analysis of linguistic features of human-generated and ChatGPT-generated texts. Dataset structure Features were calculated using… See the full description on the dataset page: https://huggingface.co/datasets/julia-lukasiewicz-pater/small-GPT-wiki-intro-features.tabulartext-classification100K<n<1M0 likes33 downloads3y agoHugging Face27JaehyungKim /p2c_polite_wikitabular1K<n<10K0 likes26 downloads3y agoHugging Face28MongoDB /wikipedia-22-12-en-voyage-embedtabular100K<n<1M0 likes25 downloads2y agoHugging Face29aisuko /simple_english_wikipedia_p0Only for the researching usage. The converting process below. # Setting the env os.environ['DATASET_URL']='http://sbert.net/datasets/simplewiki-2020-11-01.jsonl.gz' os.environ['MODEL_NAME']='multi-qa-MiniLM-L6-cos-v1' # Loading the dataset import json import gzip from sentence_transformers.util import http_get http_get(os.getenv('DATASET_URL'), os.getenv('DATASET_NAME')) passages=[] with gzip.open(os.getenv('DATASET_NAME'), 'rt', encoding='utf8') as fIn: for line in fIn:… See the full description on the dataset page: https://huggingface.co/datasets/aisuko/simple_english_wikipedia_p0.tabular100K<n<1M0 likes24 downloads3y agoHugging Face30research-dump /wiki-stance-entabular100K<n<1M0 likes24 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.