CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01softcatala /wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons. License identifiers are normalized to cc-zero, cc-by-4.0, cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self. This provides a richer alternative to Common Voice. Characteristics of the dataset: One or multiple speakers Different accents Different domain texts 761 audio files We found this dataset useful for audio tasks such as: Language detection Evaluation of STT systems New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.audioautomatic-speech-recognitionn<1K0 likes2.3k downloads2mo agoHugging Face02Intelligent-Internet /wikipedia_en wikipedia_en This is a curated Wikipedia English dataset for use with the II-Commons project. Dataset Details Dataset Description This dataset comprises a curated Wikipedia English pages. Data sourced directly from the official English Wikipedia database dump. We extract the pages, chunk them into smaller pieces, and embed them using Snowflake/snowflake-arctic-embed-m-v2.0. All vector embeddings are 16-bit half-precision vectors optimized for cosine indexing… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/wikipedia_en.tabularfeature-extraction10M<n<100M2 likes1.3k downloads1y agoHugging Face03b-mc2 /wikihow_lists Dataset Card for WikiHow Lists Dataset Summary Contains CSV of a subset of WikiHow articles. Subsets include articles that have summaries in numbered list format, unordered list of ingredients, or unordered list of items needed for the article. CSV contains a pageId to reference back to the source, title of the article, result with the list data, and a column specifying the result type (ingredient, needed items, summary) Licensing Information Data is from… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/wikihow_lists.textsummarization10K<n<100K12 likes901 downloads4y agoHugging Face04EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1007 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007.tabular10K<n<100K0 likes883 downloads24d agoHugging Face05EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1004 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1004 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1004.tabular10K<n<100K0 likes849 downloads24d agoHugging Face06vtasca /wikipedia-pageviews Wikipedia Article Pageviews This repository automatically fetches and aggregates the 100 most popular Wikipedia articles by pageviews - creating a dataset that enables tracking trending topics on Wikipedia. It works by polling the WikiMedia API on a daily basis and fetching the top 100 most popular articles from two days ago. The fetcher runs in a scheduled GitHub Actions workflow, which is available here. The dataset begins in the year 2016 and the textual data is presented as… See the full description on the dataset page: https://huggingface.co/datasets/vtasca/wikipedia-pageviews.tabularfeature-extraction100K<n<1M1 likes844 downloads4h agoHugging Face07EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1006 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006.tabular10K<n<100K0 likes842 downloads24d agoHugging Face08EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1008 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1008 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1008.tabular10K<n<100K0 likes832 downloads24d agoHugging Face09ibm-research /Wikipedia_contradict_benchmark Wikipedia contradict benchmark Wikipedia contradict benchmark is a dataset consisting of 253 high-quality, human-annotated instances designed to assess LLM performance when augmented with retrieved passages containing real-world knowledge conflicts. The dataset was created intentionally with that task in mind, focusing on a benchmark consisting of high-quality, human-annotated instances. Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/Wikipedia_contradict_benchmark.textquestion-answeringn<1K28 likes774 downloads2y agoHugging Face10vishnupriyavr /wiki-movie-plots-with-summaries Dataset Card for Wikipedia Movie Plots with AI Plot Summaries Dataset Summary Context Wikipedia Movies Plots dataset by JustinR ( https://www.kaggle.com/jrobischon/wikipedia-movie-plots ) Content Everything is the same as in https://www.kaggle.com/jrobischon/wikipedia-movie-plots Acknowledgements Please, go upvote https://www.kaggle.com/jrobischon/wikipedia-movie-plots dataset, since this is 100% based on that. Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/vishnupriyavr/wiki-movie-plots-with-summaries.text10K<n<100K6 likes738 downloads3y agoHugging Face11blo05 /cleaned_wiki_en_60-80text1M<n<10M1 likes557 downloads4y agoHugging Face12blo05 /cleaned_wiki_en_40-60text1M<n<10M1 likes503 downloads4y agoHugging Face13blo05 /cleaned_wiki_en_80-100text1M<n<10M0 likes484 downloads4y agoHugging Face14Metin /WikiRAG-TR Dataset Summary WikiRAG-TR is a dataset of 6K (5999) question and answer pairs which synthetically created from introduction part of Turkish Wikipedia Articles. The dataset is created to be used for Turkish Retrieval-Augmented Generation (RAG) tasks. Dataset Information Number of Instances: 5999 (5725 synthetically generated question-answer pairs, 274 augmented negative samples) Dataset Size: 20.5 MB Language: Turkish Dataset License: apache-2.0 Dataset Category:… See the full description on the dataset page: https://huggingface.co/datasets/Metin/WikiRAG-TR.tabularquestion-answering1K<n<10K36 likes477 downloads2y agoHugging Face15Alvaro8gb /enfermedades-wiki-marzo-2024 English Version This dataset contains detailed information on a total of 945 diseases, extracted from Wikipedia in Spanish (https://es.wikipedia.org/) in March 2024. The main purpose of this dataset is to serve as a comprehensive resource for training Large Language Models (LLMs) in Spanish, specifically for instruction tuning, pre-training, and other natural language processing (NLP) tasks. This dataset promises to be a valuable tool for research and development in Spanish language… See the full description on the dataset page: https://huggingface.co/datasets/Alvaro8gb/enfermedades-wiki-marzo-2024.texttext-generationn<1K1 likes471 downloads2y agoHugging Face16google /WikiProfile WikiProfile WikiProfile is a factual knowledge benchmark for evaluating how well language models encode and recall factual knowledge. It comprises 2,150 facts, each paired with 10 questions, for a total of 21,500 question instances. Each fact is grounded in the first paragraph (summary) of an English Wikipedia page and is defined as a proposition between two entities, a subject and an object (e.g., "Oasis played their first gig at the Boardwalk club" → subject: Oasis, object:… See the full description on the dataset page: https://huggingface.co/datasets/google/WikiProfile.tabularquestion-answering1K<n<10K20 likes466 downloads3mo agoHugging Face17wikilee /ADFA_Mappingtabular10M<n<100M0 likes460 downloads5y agoHugging Face18allegrolab /passages_wikipediatext1K<n<10K0 likes426 downloads1y agoHugging Face19EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1005 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1005 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1005.tabular10K<n<100K0 likes418 downloads24d agoHugging Face20blo05 /cleaned_wiki_enCleaned wikipedia dataset text1M<n<10M4 likes416 downloads4y agoHugging Face21Lajvi /wikihowAlltext100K<n<1M0 likes307 downloads1y agoHugging Face22blo05 /cleaned_wiki_en_20-40text1M<n<10M1 likes275 downloads4y agoHugging Face23mgmacleod /wikidata1 Wikipedia N-Link Basins A novel graph-theoretic analysis of Wikipedia's internal link structure, revealing deterministic "basins of attraction" under N-link traversal rules. Dataset Description This dataset demonstrates that Wikipedia's 17.9 million pages partition into coherent regions when following a simple rule: from any page, always follow the Nth link. Every page eventually reaches a cycle, and pages sharing the same terminal cycle form a basin of attraction.… See the full description on the dataset page: https://huggingface.co/datasets/mgmacleod/wikidata1.tabulargraph-mln<1K0 likes260 downloads9mo agoHugging Face24lewoniewski /wikipedia_quality_wikirankDatasets with WikiRank quality score as of 1 August 2024 for 47 million Wikipedia articles in 55 language versions (also simplified version for each language in separate files). The WikiRank quality score is a metric designed to assess the overall quality of a Wikipedia article. Although its specific algorithm can vary depending on the implementation, the score typically combines several key features of the Wikipedia article. Why It’s Important Enhances Trust: For readers and… See the full description on the dataset page: https://huggingface.co/datasets/lewoniewski/wikipedia_quality_wikirank.tabular10M<n<100M4 likes223 downloads2y agoHugging Face25anhaidgroup /polaris-wikitables-v2 WikiTables v2 WikiTables is one of six datasets in Polaris: Learning to Generate Table Descriptions from Retrieval Feedback, alongside aw, arctic, lter, ecir, and wtr. It holds 3,361 tables scraped from Wikipedia articles and 57 keyword queries over them. For each query–table pair, a person scored how well that table answers that query; those scores are the relevance judgments, and they live in qrels.csv. The tables have no names. What describes a table is its column names and… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-wikitables-v2.tabulartext-retrieval1K<n<10K0 likes219 downloads1mo agoHugging Face26MongoDB /cosmopedia-wikihow-chunked Overview This dataset is a chunked version of a subset of data in the Cosmopedia dataset curated by Hugging Face. Specifically, we have only used a subset of Wikihow articles from the Cosmopedia dataset, and each article has been split into chunks containing no more than 2 paragraphs. Dataset Structure Each record in the dataset represents a chunk of a larger article, and contains the following fields: doc_id: A unique identifier for the parent article chunk_id: A unique… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/cosmopedia-wikihow-chunked.tabularquestion-answering1M<n<10M9 likes204 downloads3y agoHugging Face27aadityaubhat /GPT-wiki-intro GPT Wiki Intro Overview Dataset for training models to classify human written vs GPT/ChatGPT generated text. This dataset contains Wikipedia introductions and GPT (Curie) generated introductions for 150k topics. Prompt used for generating text 200 word wikipedia style introduction on '{title}' {starter_text} where title is the title for the wikipedia page, and starter_text is the first seven words of the wikipedia introduction. Here's an example of prompt used to… See the full description on the dataset page: https://huggingface.co/datasets/aadityaubhat/GPT-wiki-intro.tabulartext-classification100K<n<1M27 likes173 downloads3y agoHugging Face28yashassnadig /wikimovies Wikipedia Movies Dataset Dataset Description This dataset contains 58.1k movie information scraped from Wikipedia's "List of films" pages. The data includes basic movie metadata, infobox information, and introductory text from individual movie Wikipedia pages. NOTE: This is an uncleaned dataset containing raw scraped data. The content is sourced from Wikipedia and is not owned by the dataset creator. All content remains under Wikipedia's licensing terms.… See the full description on the dataset page: https://huggingface.co/datasets/yashassnadig/wikimovies.tabulartext-classification10K<n<100K1 likes173 downloads1y agoHugging Face29maywell /ko_wikidata_QA 업데이트 로그 2023-11-03 : MarkrAI의 Dedup 적용. 한국어 위키 데이터 QA셋 본 데이터는 Synatra-7B-Instruct 모델과 ChatGPT를 사용하여, 제작된 QA셋입니다. 해당 데이터를 직접적으로 상업적으로 사용하는 것은 허용되지 않으며, 데이터를 이용하여 훈련된 모델에 대한 상업적 사용은 허용됩니다. 아직 완벽히 정제되지는 않았으며, 오류나 수정사항에 대해서는 PR 부탁드립니다. text100K<n<1M43 likes170 downloads3y agoHugging Face30blo05 /wiki_titlestext1M<n<10M0 likes165 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.