datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikipedia_ruMoeGirlPedia_wikitext_raw_archiveGlad to see models and datasets were inspired from this dataset, thanks to all who are using this dataset in their training materials.
Feel free to re-upload the contents to places like the Internet Archive (Please follow the license and keep these files as-is) to help preserve this digital asset.
Looking forward to see more models and synthetic datasets trained from this raw archive, good luck!
Note: Due to the content censorship system introduced by MGP on 2024/03/29, it is unclear that… See the full description on the dataset page: https://huggingface.co/datasets/milashkaarshif/MoeGirlPedia_wikitext_raw_archive.wikipedia_envel_commons_wikidata
Visual Entity Linking: Wikimedia Commons & Wikidata
This dataset allows to train and evaluate ML models that link Wikimedia Commons images to the Wikidata items they depict.
Disclaimer: All images contained in this dataset are generally assumed to be freely usable (as intended for Wikimedia Commons). Each image's license and author/
uploader is - to the best of our ability - reported in its metadata (see section Dataset Structure). If you want your image's attribution changed or the… See the full description on the dataset page: https://huggingface.co/datasets/aiintelligentsystems/vel_commons_wikidata.wikiart_with_BLIP_captionsWikipedia-Knowledge-2M
📃 Paper | 🤗 Hugging Face | ⭐ Github
Dataset Overview
In the table below, we provide a brief summary of the dataset statistics.
Category
Size
Total Sample
2019163
Total Image
2019163
Average Answer Length
84
Maximum Answer Length
5851
JSON Overview
Each dictionary in the JSON file contains three keys: 'id', 'image', and 'conversations'.
The 'id' is the unique identifier for the current data in the entire dataset.
The 'image' stores… See the full description on the dataset page: https://huggingface.co/datasets/Ghaser/Wikipedia-Knowledge-2M.wikivideo
Paper and Code
Associated with the paper: WikiVideo (https://arxiv.org/abs/2504.00939)
Associated with the github repository (https://github.com/alexmartin1722/wikivideo)
Download instructions
The dataset can be found on huggingface. However, you can't use the datasets library to access the videos because everything is tarred. Instead you need to locally download the dataset and then untar the videos (and audios if you use those).
Step 1: Install git-lfs
The… See the full description on the dataset page: https://huggingface.co/datasets/hltcoe/wikivideo.WikiArt-81K-BLIP_2-768x768
WikiArt Resized Dataset
Description
This dataset contains 81,444 artistic images from WikiArt, organized into different artistic genres. The images have been resized to a uniform resolution of 768x768 pixels using LANCZOS resampling and padded to maintain aspect ratio, ensuring consistency for machine learning tasks and computational art analysis. The base for the dataset was Dant33/WikiArt-81K-BLIP_2-captions.
Enhancements
1. Image Resizing… See the full description on the dataset page: https://huggingface.co/datasets/Dant33/WikiArt-81K-BLIP_2-768x768.Wikiart_with_StableDiffusion
Artistic Images Transformed by Stable Diffusion XL Refiner 1.0
Overview
This dataset contains 81,444 AI-generated images derived from famous paintings across 27 artistic genres. The transformation process involved resizing the original images to 768px, generating detailed descriptions using BLIP2, and creating customized prompts with LLaMA 3 8B. These prompts were then used with Stable Diffusion XL Refiner 1.0 to generate modified versions of the original artworks.
The… See the full description on the dataset page: https://huggingface.co/datasets/Dant33/Wikiart_with_StableDiffusion.WikipediaDumpWikiArt-81K-BLIP_2-1024x1024
WikiArt Resized Dataset
Description
This dataset contains 81,444 artistic images from WikiArt, organized into different artistic genres. The images have been resized to a uniform resolution of 1024x1024 pixels using LANCZOS resampling, ensuring consistency for machine learning tasks and computational art analysis. The base for the dataset was Dant33/WikiArt-81K-BLIP_2-captions
Enhancements
1. Image Resizing
All images have been resized to… See the full description on the dataset page: https://huggingface.co/datasets/Dant33/WikiArt-81K-BLIP_2-1024x1024.WikiArt-81K-BLIP_2-captions
WikiArt Enhanced Dataset
Description
This dataset contains 81,444 artistic images from WikiArt, organized into different artistic genres. It has undergone several improvements and corrections to optimize its use in machine learning tasks and computational art analysis. Credits to the original author of daset go to: WikiArt
Enhancements
1. Encoding Issues Correction
Fixed encoding issues in filenames and artist information.
All filenames were renamed… See the full description on the dataset page: https://huggingface.co/datasets/Dant33/WikiArt-81K-BLIP_2-captions.openslr_enhanced0504_combination_instruction_wikihowwikilargeimgwikisumemotional-tts-wikiwikibig
WikiBig
Large dataset made of chunks of images found on https://commons.wikimedia.org/wiki/Category:Large_images.
Each chunk has 256x256 pixels.
WikipediaDumpwikicommons-cc-0Wikivideo_combined_video
