racineai/VDR_MEGA_2
VDR_MEGA_2 Dataset Summary VDR_MEGA_2 is a high-quality multimodal dataset created through the merge of multiple domain-specific datasets with enhanced data processing techniques. This dataset represents our most refined approach to multimodal data generation, incorporating filtering algorithms and improved AI-assisted content generation to deliver superior quality for RAG, DSE, question answering, document search, and vision-language model training tasks.… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_2.
VDRMEGA2
Dataset Summary
VDR_MEGA_2 is a high-quality multimodal dataset created through the merge of multiple domain-specific datasets with enhanced data processing techniques. This dataset represents our most refined approach to multimodal data generation, incorporating filtering algorithms and improved AI-assisted content generation to deliver superior quality for RAG, DSE, question answering, document search, and vision-language model training tasks.
Source Datasets
This merged dataset combines the the following datasets:
Data Fields
Each entry contains:
- `id` (string): Unique identifier
- `query` (string): High-quality technical/domain-specific question
- `image` (PIL.Image): High-resolution visual rendering of source document page
- `language` (string): Detected language of the image (queries sometimes differ on purpose)
Dataset Curators
- Léo Appourchaux
- Paul Lemaistre
- Yumeng Ye
- Mattéo KHAN
- André-Louis Rochet
