datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-DatasetTunisian Proverbs with Image Associations: A Cultural and Linguistic Dataset
Description
This dataset explores the rich oral tradition of Tunisian proverbs mapped into text format, pairing each with contextual explanations, English translations both word-to-word and it's equivalent Target Language dynamic, Automated prompt and AI-generated visual interpretations.
It bridges linguistic, cultural, and visual modalities making it valuable for tasks in cross-cultural NLP, generative… See the full description on the dataset page: https://huggingface.co/datasets/HabibaAbderrahim/Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-Dataset.Argimi-Ardian-Finance-10k-text-image
The ArGiMI Ardian datasets : text and images
The ArGiMi project is committed to open-source principles and data sharing.
Thanks to our generous partners, we are releasing several valuable datasets to the public.
Dataset description
This dataset comprises 34,000 financial annual reports, written in English, meticulously
extracted from their original PDF format to provide a valuable resource for researchers and developers in financial
analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.62k-images-khmer-printed-dataset
62k Khmer-English Printed Dataset
This repository contains a dataset of Khmer and English printed text images for training, validation, and testing. The dataset is stored in parquet format and managed using Git Large File Storage (LFS).
Installation
Prerequisites
Before cloning this repository, make sure you have Git LFS installed:
Install Git LFS
Linux/macOS:curl -s https://packagecloud.io/install/repositories/github/git-lfs/script.deb.sh | sudo… See the full description on the dataset page: https://huggingface.co/datasets/SoyVitou/62k-images-khmer-printed-dataset.astrobridge-image-captions
AstroBridge Legacy Survey Captions
3,487 imaging cutouts from the Legacy Survey (DR10 South + North), crossmatched against
published literature mentions and captioned in four independent stages by Gemini
(gemini-3.7-flash), following the AstroLLaVA data-generation approach (Zaman et al. 2025,
arXiv:2504.08583): no caption is ever told the object's
real name or catalog designation, and no caption states a fact that isn't derivable from the
pixels or the (redacted-at-the-model… See the full description on the dataset page: https://huggingface.co/datasets/gapatron/astrobridge-image-captions.LaTeX_Image_Pairs
LaTeX Image Pairs Dataset
This dataset comprises a unique collection of LaTeX expressions paired with their corresponding images. The LaTeX expressions were meticulously scraped from a variety of open-source textbooks, ensuring a diverse and comprehensive dataset. Sample references from these textbooks will be provided to illustrate the sources of these expressions.
In addition to the raw LaTeX expressions, this dataset includes images of the rendered expressions. Each LaTeX… See the full description on the dataset page: https://huggingface.co/datasets/henryholloway/LaTeX_Image_Pairs.bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bghira/pseudo-camera-10k with responses/captions generated with gemini-2.0-flash-thinking-exp-1219.
The format should be similar to that of liuhaotian/LLaVA-Instruct-150K.
Images can be found in the images.zip folder. The zip also contains .txt captions for ease of use in non-VQA tasks.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Images/bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.recipe-synthetic-images-10k
Recipe PDF Dataset
A multimodal dataset of 10K+ recipes rendered as PDF images with full metadata.
Dataset Description
Each sample contains:
image: Recipe rendered as a styled PDF page (PNG, ~1654x2339px)
name: Recipe title
description: Recipe description
ingredients: List of ingredients
steps: Cooking instructions
nutrition: Nutritional values (calories, fat%, sugar%, sodium%, protein%, sat.fat%, carbs%)
random_reviews: User reviews
minutes: Cooking time
tags: Recipe… See the full description on the dataset page: https://huggingface.co/datasets/TurkishCodeMan/recipe-synthetic-images-10k.Handwritten-Historical-Archive-Image-Dataset-of-Modern-China_1840_1949
📜 Chinese Modern Era (1840–1949) Handwritten Historical Archive Dataset
中国近代史 (1840–1949) 手写历史档案数据集
📖 Dataset Description | 数据集描述
🎯 Purpose & Motivation | 目的与动机
To address the recognition difficulties and generalization bottlenecks faced by existing Optical Character Recognition (OCR) models when processing handwritten historical archives from modern Chinese history (1840–1949), a joint student research team from Capital Normal University… See the full description on the dataset page: https://huggingface.co/datasets/JIA244601/Handwritten-Historical-Archive-Image-Dataset-of-Modern-China_1840_1949.Chinese-Image-Text-Corpus-dataset
REILX/Chinese-Image-Text-Corpus-dataset
[ English | 中文 ]
Introduction
The REILX/Chinese-Image-Text-Corpus-dataset is a multimodal dataset that pairs Chinese textual data with corresponding images. This dataset is derived from the Chinese-Xinhua Dictionary Database, which includes idioms, single characters, words, and aphorisms.
Dataset Structure
The dataset is organized into the following categories:
Idioms: Traditional Chinese idioms with explanations and… See the full description on the dataset page: https://huggingface.co/datasets/REILX/Chinese-Image-Text-Corpus-dataset.Handpicked-Images-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
Handpicked-Images-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
Some random images with responses/captions generated with gemini-2.0-flash-thinking-exp-1219.
The format should be similar to that of liuhaotian/LLaVA-Instruct-150K.
Images can be found in the images.zip folder. The zip also contains .txt captions for ease of use in non-VQA tasks.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Images/Handpicked-Images-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.kenya-bee-health-qa-image-triples
Kenya Bee Health Training Data
This folder is a starter database for BeeCare Anywhere / Gemma Apiary. It is intentionally small, transparent, and license-aware: use it to prove the Q/A/image-triple pipeline, then expand it with Kenyan field data before trusting model behavior in production.
Important Model Note
google/gemma-2b is a text-to-text, decoder-only model. It cannot directly read pictures. Use these image triples with a vision-capable model path, for… See the full description on the dataset page: https://huggingface.co/datasets/yahelr1/kenya-bee-health-qa-image-triples.CuPer_Images
CuPer_Images
Resumen del Dataset
CuPer_Images es un dataset multimodal especializado en culturas precolombinas de Perú, desarrollado por NovaIA, el laboratorio de inteligencia artificial de Grupo Neura. Este dataset fue utilizado como parte del entrenamiento de Amaru, un modelo de lenguaje de propósito general con conocimientos profundos en las civilizaciones ancestrales del Perú.
Información del Dataset
Tamaño total: 4,247 muestras
División: 3,396 muestras… See the full description on the dataset page: https://huggingface.co/datasets/NovaIALATAM/CuPer_Images.finance-legal-mrc-with-images
🧾 finance-legal-mrc-with-images (tableqa/test)
Multimodal-ready table image dataset designed for TIG (Table Information Generation) inputTIG 입력 전용 테이블 이미지 데이터셋 (VLM 활용 가능)
📦 Dataset Overview | 데이터셋 개요
Feature
Description (EN)
설명 (KR)
🧩 Split
tableqa / test
tableqa / test 스플릿
📄 Total Rows
1,197 (unique tables only)
총 1,197건 (중복 제거된 고유 테이블 기준)
🖼️ Image Format
PNG (rendered from raw HTML tables)
원본 HTML 테이블을 PNG 이미지로 렌더링
🔗 Source
Derived from… See the full description on the dataset page: https://huggingface.co/datasets/didi0di/finance-legal-mrc-with-images.coco_image_extract
Modified Coco Dataset Files
Required dependencies
OpenCV (cv2):
pip install opencv-python
img_data.psv
Extract of the coco dataset containing the following labels: ["airplane", "backpack", "cell phone", "handbag", "suitcase", "knife", "laptop", "car"]
Structured as follows:
| Field | Description |
| --------------- |… See the full description on the dataset page: https://huggingface.co/datasets/iix/coco_image_extract.ImageText-Question-answer-pairs-58K-Claude-3.5-Sonnnet
REILX/ImageText-Question-answer-pairs-58K-Claude-3.5-Sonnnet
从VisualGenome数据集V1.2中随机抽取21717张图片,利用Claude-3-opus-20240229和Claude-3-sonnet-20240620两个模型生成了总计58312个问答对,每张图片约3个问答,其中必有一个关于图像细节的问答。Claude-3-opus-20240229模型贡献了约3,028个问答对,而Claude-3-sonnet-20240620模型则生成了剩余的问答对。
Code
使用以下代码生成问答对:
# -*- coding: gbk -*-
import os
import random
import shutil
import re
import json
import requests
import base64
import time
from tqdm import tqdm
from json_repair import repair_json… See the full description on the dataset page: https://huggingface.co/datasets/REILX/ImageText-Question-answer-pairs-58K-Claude-3.5-Sonnnet.
