CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Alvaro8gb /enfermedades-wiki-marzo-2024 English Version This dataset contains detailed information on a total of 945 diseases, extracted from Wikipedia in Spanish (https://es.wikipedia.org/) in March 2024. The main purpose of this dataset is to serve as a comprehensive resource for training Large Language Models (LLMs) in Spanish, specifically for instruction tuning, pre-training, and other natural language processing (NLP) tasks. This dataset promises to be a valuable tool for research and development in Spanish language… See the full description on the dataset page: https://huggingface.co/datasets/Alvaro8gb/enfermedades-wiki-marzo-2024.texttext-generationn<1K1 likes447 downloads3y agoHugging Face02aadityaubhat /GPT-wiki-intro GPT Wiki Intro Overview Dataset for training models to classify human written vs GPT/ChatGPT generated text. This dataset contains Wikipedia introductions and GPT (Curie) generated introductions for 150k topics. Prompt used for generating text 200 word wikipedia style introduction on '{title}' {starter_text} where title is the title for the wikipedia page, and starter_text is the first seven words of the wikipedia introduction. Here's an example of prompt used to… See the full description on the dataset page: https://huggingface.co/datasets/aadityaubhat/GPT-wiki-intro.tabulartext-classification100K<n<1M27 likes198 downloads3y agoHugging Face03yashassnadig /wikimovies Wikipedia Movies Dataset Dataset Description This dataset contains 58.1k movie information scraped from Wikipedia's "List of films" pages. The data includes basic movie metadata, infobox information, and introductory text from individual movie Wikipedia pages. NOTE: This is an uncleaned dataset containing raw scraped data. The content is sourced from Wikipedia and is not owned by the dataset creator. All content remains under Wikipedia's licensing terms.… See the full description on the dataset page: https://huggingface.co/datasets/yashassnadig/wikimovies.tabulartext-classification10K<n<100K1 likes176 downloads1y agoHugging Face04ZDZR /BRINK-Wikidata5m BRINK-Wikidata5m BRINK (Benchmark for Reasoning under Incomplete Knowledge) is a benchmark for evaluating Knowledge Graph–based Retrieval-Augmented Generation (KG-RAG) under incomplete knowledge. Unlike standard KGQA benchmarks, BRINK is designed so that each question cannot be answered by directly retrieving a single explicit supporting triple. Instead, the answer must be inferred from alternative reasoning paths that remain in the graph after the directly supporting fact is… See the full description on the dataset page: https://huggingface.co/datasets/ZDZR/BRINK-Wikidata5m.textquestion-answering10K<n<100K1 likes101 downloads6mo agoHugging Face05muset-ai /Wiki_Live_Challenge Wiki Live Challenge Dataset [English | 中文] English 📖 Dataset Overview This is the official dataset accompanying the Wiki Live Challenge benchmark. It contains Wikipedia Good Articles (GAs) as ground truth and research articles generated by leading deep research AI systems. Wiki Live Challenge is the first live benchmark for evaluating Deep Research Agents (DRAs) on their ability to generate Wikipedia-quality articles. Unlike static benchmarks, Wiki Live… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/Wiki_Live_Challenge.tabulartext-generationn<1K1 likes72 downloads8mo agoHugging Face06Svngoku /wikipedia-2023-11-kikongo-lingala-cohere-multilingual-v3texttext-generation10K<n<100K2 likes70 downloads2y agoHugging Face07gozh /habr_and_wikipedia1gb Russian-English dataset containing articles from Habr and Wikipedia. texttext-generation100K<n<1M1 likes64 downloads3y agoHugging Face08BrightData /Wikipedia-Articles Dataset Card for "BrightData/Wikipedia-Articles" Dataset Summary Explore a collection of millions of Wikipedia articles with the Wikipedia dataset, comprising over 1.23M structured records and 10 data fields updated and refreshed regularly. Each entry includes all major data points such as timestamp, URLs, article titles, raw and cataloged text, images, "see also" references, external links, and a structured table of contents. For a complete list of data points, please… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/Wikipedia-Articles.texttext-classification100K<n<1M7 likes56 downloads2y agoHugging Face09orionai /en_wikipedia_001 Dataset Card for en_wikipedia_001 The en_wikipedia_001 dataset is a collection of crawled paragraph text from Wikipedia on the 28th of April, 2024. It contains high-quality text, stored in multiple documents, available to be used to finetune or train AI models based that the license is followed. Dataset Details The dataset was crawled using our web crawler on the 28th of April at an average of 1 page per second as to respect robots.txt rules. Strict licensing must be… See the full description on the dataset page: https://huggingface.co/datasets/orionai/en_wikipedia_001.textquestion-answeringn<1K2 likes34 downloads2y agoHugging Face10leks-forever /lez_wiki_20240920texttext-generation1K<n<10K1 likes15 downloads2y agoHugging Face11cantonesesra /Cantonese_Wiki_DoYouKnow_1.5K Cantonese Question Dataset from Yue Wiki A collection of questions in Cantonese, extracted from Yue Wiki. This dataset contains a variety of questions covering different topics and domains. Disclaimer The content and opinions expressed in this dataset do not represent the views, beliefs, or positions of the dataset creators, contributors, or hosting organizations. This dataset is provided solely for the purpose of improving AI systems' understanding of the Cantonese… See the full description on the dataset page: https://huggingface.co/datasets/cantonesesra/Cantonese_Wiki_DoYouKnow_1.5K.texttext-generation1K<n<10K0 likes15 downloads1y agoHugging Face12Lam-ia /Wikipedia-Euskeratexttext-generationn<1K0 likes11 downloads3y agoHugging Face13Navanjana /Sinhala-Wikitexttext-generation10K<n<100K0 likes9 downloads2y agoHugging Face14Lizagrin /wikiart_captions WikiArt Captions Subset — Multimodal Art Retrieval Dataset This dataset is a curated subset of 6,000 paintings from the WikiArt collection.It was created as part of a project on multimodal art retrieval, combining visual, textual, and semantic information. Each record represents one artwork and includes: Field Description image_row Row index in the source subset (integer) caption Automatically generated textual description (caption) using the BLIP model… See the full description on the dataset page: https://huggingface.co/datasets/Lizagrin/wikiart_captions.texttext-generation1K<n<10K1 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.