datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikipedia_ruwikipedia_enWikipedia-Knowledge-2M
📃 Paper | 🤗 Hugging Face | ⭐ Github
Dataset Overview
In the table below, we provide a brief summary of the dataset statistics.
Category
Size
Total Sample
2019163
Total Image
2019163
Average Answer Length
84
Maximum Answer Length
5851
JSON Overview
Each dictionary in the JSON file contains three keys: 'id', 'image', and 'conversations'.
The 'id' is the unique identifier for the current data in the entire dataset.
The 'image' stores… See the full description on the dataset page: https://huggingface.co/datasets/Ghaser/Wikipedia-Knowledge-2M.WikipediaDumpWikipediaDump
