datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
enfermedades-wiki-marzo-2024
English Version
This dataset contains detailed information on a total of 945 diseases, extracted from Wikipedia in Spanish (https://es.wikipedia.org/) in March 2024. The main purpose of this dataset is to serve as a comprehensive resource for training Large Language Models (LLMs) in Spanish, specifically for instruction tuning, pre-training, and other natural language processing (NLP) tasks. This dataset promises to be a valuable tool for research and development in Spanish language… See the full description on the dataset page: https://huggingface.co/datasets/Alvaro8gb/enfermedades-wiki-marzo-2024.GPT-wiki-intro
GPT Wiki Intro
Overview
Dataset for training models to classify human written vs GPT/ChatGPT generated text.
This dataset contains Wikipedia introductions and GPT (Curie) generated introductions for 150k topics.
Prompt used for generating text
200 word wikipedia style introduction on '{title}'
{starter_text}
where title is the title for the wikipedia page, and starter_text is the first seven words of the wikipedia introduction.
Here's an example of prompt used to… See the full description on the dataset page: https://huggingface.co/datasets/aadityaubhat/GPT-wiki-intro.wikimovies
Wikipedia Movies Dataset
Dataset Description
This dataset contains 58.1k movie information scraped from Wikipedia's "List of films" pages. The data includes basic movie metadata, infobox information, and introductory text from individual movie Wikipedia pages.
NOTE: This is an uncleaned dataset containing raw scraped data. The content is sourced from Wikipedia and is not owned by the dataset creator. All content remains under Wikipedia's licensing terms.… See the full description on the dataset page: https://huggingface.co/datasets/yashassnadig/wikimovies.BRINK-Wikidata5m
BRINK-Wikidata5m
BRINK (Benchmark for Reasoning under Incomplete Knowledge) is a benchmark for evaluating Knowledge Graph–based Retrieval-Augmented Generation (KG-RAG) under incomplete knowledge. Unlike standard KGQA benchmarks, BRINK is designed so that each question cannot be answered by directly retrieving a single explicit supporting triple. Instead, the answer must be inferred from alternative reasoning paths that remain in the graph after the directly supporting fact is… See the full description on the dataset page: https://huggingface.co/datasets/ZDZR/BRINK-Wikidata5m.Wiki_Live_Challenge
Wiki Live Challenge Dataset
[English | 中文]
English
📖 Dataset Overview
This is the official dataset accompanying the Wiki Live Challenge benchmark. It contains Wikipedia Good Articles (GAs) as ground truth and research articles generated by leading deep research AI systems.
Wiki Live Challenge is the first live benchmark for evaluating Deep Research Agents (DRAs) on their ability to generate Wikipedia-quality articles. Unlike static benchmarks, Wiki Live… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/Wiki_Live_Challenge.wikipedia-2023-11-kikongo-lingala-cohere-multilingual-v3habr_and_wikipedia1gb Russian-English dataset containing articles from Habr and Wikipedia.
Wikipedia-Articles
Dataset Card for "BrightData/Wikipedia-Articles"
Dataset Summary
Explore a collection of millions of Wikipedia articles with the Wikipedia dataset, comprising over 1.23M structured records and 10 data fields updated and refreshed regularly.
Each entry includes all major data points such as timestamp, URLs, article titles, raw and cataloged text, images, "see also" references, external links, and a structured table of contents.
For a complete list of data points, please… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/Wikipedia-Articles.en_wikipedia_001
Dataset Card for en_wikipedia_001
The en_wikipedia_001 dataset is a collection of crawled paragraph text from Wikipedia on the 28th of April, 2024. It contains high-quality text, stored in multiple documents, available to be used to finetune or train AI models based that the license is followed.
Dataset Details
The dataset was crawled using our web crawler on the 28th of April at an average of 1 page per second as to respect robots.txt rules. Strict licensing must be… See the full description on the dataset page: https://huggingface.co/datasets/orionai/en_wikipedia_001.lez_wiki_20240920Cantonese_Wiki_DoYouKnow_1.5K
Cantonese Question Dataset from Yue Wiki
A collection of questions in Cantonese, extracted from Yue Wiki. This dataset contains a variety of questions covering different topics and domains.
Disclaimer
The content and opinions expressed in this dataset do not represent the views, beliefs, or positions of the dataset creators, contributors, or hosting organizations. This dataset is provided solely for the purpose of improving AI systems' understanding of the Cantonese… See the full description on the dataset page: https://huggingface.co/datasets/cantonesesra/Cantonese_Wiki_DoYouKnow_1.5K.Wikipedia-EuskeraSinhala-Wikiwikiart_captions
WikiArt Captions Subset — Multimodal Art Retrieval Dataset
This dataset is a curated subset of 6,000 paintings from the WikiArt collection.It was created as part of a project on multimodal art retrieval, combining visual, textual, and semantic information.
Each record represents one artwork and includes:
Field
Description
image_row
Row index in the source subset (integer)
caption
Automatically generated textual description (caption) using the BLIP model… See the full description on the dataset page: https://huggingface.co/datasets/Lizagrin/wikiart_captions.
