datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-DatasetTunisian Proverbs with Image Associations: A Cultural and Linguistic Dataset
Description
This dataset explores the rich oral tradition of Tunisian proverbs mapped into text format, pairing each with contextual explanations, English translations both word-to-word and it's equivalent Target Language dynamic, Automated prompt and AI-generated visual interpretations.
It bridges linguistic, cultural, and visual modalities making it valuable for tasks in cross-cultural NLP, generative… See the full description on the dataset page: https://huggingface.co/datasets/HabibaAbderrahim/Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-Dataset.Argimi-Ardian-Finance-10k-text-image
The ArGiMI Ardian datasets : text and images
The ArGiMi project is committed to open-source principles and data sharing.
Thanks to our generous partners, we are releasing several valuable datasets to the public.
Dataset description
This dataset comprises 34,000 financial annual reports, written in English, meticulously
extracted from their original PDF format to provide a valuable resource for researchers and developers in financial
analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.X-RAY_images_with_reports
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/mariam-srour/X-RAY_images_with_reports.62k-images-khmer-printed-dataset
62k Khmer-English Printed Dataset
This repository contains a dataset of Khmer and English printed text images for training, validation, and testing. The dataset is stored in parquet format and managed using Git Large File Storage (LFS).
Installation
Prerequisites
Before cloning this repository, make sure you have Git LFS installed:
Install Git LFS
Linux/macOS:curl -s https://packagecloud.io/install/repositories/github/git-lfs/script.deb.sh | sudo… See the full description on the dataset page: https://huggingface.co/datasets/SoyVitou/62k-images-khmer-printed-dataset.astrobridge-image-captions
AstroBridge Legacy Survey Captions
3,487 imaging cutouts from the Legacy Survey (DR10 South + North), crossmatched against
published literature mentions and captioned in four independent stages by Gemini
(gemini-3.7-flash), following the AstroLLaVA data-generation approach (Zaman et al. 2025,
arXiv:2504.08583): no caption is ever told the object's
real name or catalog designation, and no caption states a fact that isn't derivable from the
pixels or the (redacted-at-the-model… See the full description on the dataset page: https://huggingface.co/datasets/gapatron/astrobridge-image-captions.asil-benchmark-images
ASIL Benchmark Images
This repository hosts prebuilt ASIL benchmark runtime images for the v0.1.0
paper reproduction release.
The code release is available at https://github.com/sharryXR/ASIL.
Paper: ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions
Project page: https://sharryxr.github.io/ASIL
Contents
manifest.json: image manifest with version, size, SHA256, and definition
file hashes.
benchmark_images.v0.1.0.json: versioned copy… See the full description on the dataset page: https://huggingface.co/datasets/sharryXR/asil-benchmark-images.pixmo-points-with-image
Pixmo pointing downloaded without hash checking
Please concatenate the three files. The images have been renamed using their hash keys.
The images used in this project were obtained from the publicly available links provided by the (original dataset)[https://huggingface.co/datasets/allenai/pixmo-points]. We only downloaded and organized these images for research and non-commercial purposes. All copyrights and related responsibilities remain with the original copyright holders.… See the full description on the dataset page: https://huggingface.co/datasets/ShuaiYang03/pixmo-points-with-image.NFT-70M_image
Dataset Card for "NFT-70M_image"
Dataset summary
The NFT-70M_image dataset is a companion for our released NFT-70M_transactions dataset,
which is the largest and most up-to-date collection of Non-Fungible Tokens (NFT) transactions between 2021 and 2023 sourced from OpenSea.
As we also reported in the "Data anonymization" section of the dataset card of NFT-70M_transactions,
the URLs of NFT images data were replaced by identifiers to numerical vectors that represent an… See the full description on the dataset page: https://huggingface.co/datasets/MLNTeam-Unical/NFT-70M_image.Open-Personix
Open-Personix
Dataset Summary
Open-Personix is a structured JSON dataset maintained under Poralus.
The dataset is primarily text and metadata: each record contains a relative image path,
a natural-language caption, and descriptive annotation fields for a person-centered sample.
The dataset is designed for workflows such as:
caption generation and caption analysis
text-based filtering over person annotations
metadata-aware retrieval and evaluation
multimodal experiments… See the full description on the dataset page: https://huggingface.co/datasets/Below-Image/Open-Personix.LaTeX_Image_Pairs
LaTeX Image Pairs Dataset
This dataset comprises a unique collection of LaTeX expressions paired with their corresponding images. The LaTeX expressions were meticulously scraped from a variety of open-source textbooks, ensuring a diverse and comprehensive dataset. Sample references from these textbooks will be provided to illustrate the sources of these expressions.
In addition to the raw LaTeX expressions, this dataset includes images of the rendered expressions. Each LaTeX… See the full description on the dataset page: https://huggingface.co/datasets/henryholloway/LaTeX_Image_Pairs.DLSite-Meta-And-Sample-Images-860k
DLSite-Meta-And-Sample-Images-860k
This dataset contains almost all metadata and links to sample images from DLSite.
BLUEX_without_images
BLUEX: A benchmark based on Brazilian Leading Universities Entrance eXams
BLUEX is a multimodal dataset consisting of the two leading university entrance exams conducted in Brazil: Convest (Unicamp) and Fuvest (USP), spanning from 2018 to 2024. The benchmark comprises 724 questions that do not have accompanying images. This task evaluates models on the multiple-choice questions from these exams. The model is given a question in Portuguese together with its answer choices and… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/BLUEX_without_images.recipe-synthetic-images-10k
Recipe PDF Dataset
A multimodal dataset of 10K+ recipes rendered as PDF images with full metadata.
Dataset Description
Each sample contains:
image: Recipe rendered as a styled PDF page (PNG, ~1654x2339px)
name: Recipe title
description: Recipe description
ingredients: List of ingredients
steps: Cooking instructions
nutrition: Nutritional values (calories, fat%, sugar%, sodium%, protein%, sat.fat%, carbs%)
random_reviews: User reviews
minutes: Cooking time
tags: Recipe… See the full description on the dataset page: https://huggingface.co/datasets/TurkishCodeMan/recipe-synthetic-images-10k.Handwritten-Historical-Archive-Image-Dataset-of-Modern-China_1840_1949
📜 Chinese Modern Era (1840–1949) Handwritten Historical Archive Dataset
中国近代史 (1840–1949) 手写历史档案数据集
📖 Dataset Description | 数据集描述
🎯 Purpose & Motivation | 目的与动机
To address the recognition difficulties and generalization bottlenecks faced by existing Optical Character Recognition (OCR) models when processing handwritten historical archives from modern Chinese history (1840–1949), a joint student research team from Capital Normal University… See the full description on the dataset page: https://huggingface.co/datasets/JIA244601/Handwritten-Historical-Archive-Image-Dataset-of-Modern-China_1840_1949.bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bghira/pseudo-camera-10k with responses/captions generated with gemini-2.0-flash-thinking-exp-1219.
The format should be similar to that of liuhaotian/LLaVA-Instruct-150K.
Images can be found in the images.zip folder. The zip also contains .txt captions for ease of use in non-VQA tasks.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Images/bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.Chinese-Image-Text-Corpus-dataset
REILX/Chinese-Image-Text-Corpus-dataset
[ English | 中文 ]
Introduction
The REILX/Chinese-Image-Text-Corpus-dataset is a multimodal dataset that pairs Chinese textual data with corresponding images. This dataset is derived from the Chinese-Xinhua Dictionary Database, which includes idioms, single characters, words, and aphorisms.
Dataset Structure
The dataset is organized into the following categories:
Idioms: Traditional Chinese idioms with explanations and… See the full description on the dataset page: https://huggingface.co/datasets/REILX/Chinese-Image-Text-Corpus-dataset.Handpicked-Images-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
Handpicked-Images-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
Some random images with responses/captions generated with gemini-2.0-flash-thinking-exp-1219.
The format should be similar to that of liuhaotian/LLaVA-Instruct-150K.
Images can be found in the images.zip folder. The zip also contains .txt captions for ease of use in non-VQA tasks.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Images/Handpicked-Images-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.kenya-bee-health-qa-image-triples
Kenya Bee Health Training Data
This folder is a starter database for BeeCare Anywhere / Gemma Apiary. It is intentionally small, transparent, and license-aware: use it to prove the Q/A/image-triple pipeline, then expand it with Kenyan field data before trusting model behavior in production.
Important Model Note
google/gemma-2b is a text-to-text, decoder-only model. It cannot directly read pictures. Use these image triples with a vision-capable model path, for… See the full description on the dataset page: https://huggingface.co/datasets/yahelr1/kenya-bee-health-qa-image-triples.frontend-react-dataset-no-images
Frontend React Dataset — No Images
Image-free derivative of Reubencf/frontend-react-dataset.
This version preserves all 1,000 examples while completely removing the screenshot column and its embedded image bytes.
Schema
Column
Type
Description
description
string
Description or prompt for the interface
response
string
React/TSX implementation response
from datasets import load_dataset
ds = load_dataset("Reubencf/frontend-react-dataset-no-images")… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/frontend-react-dataset-no-images.dual-stream-image-prompts
Dual-Stream Image Prompts
Multi-dialect image-prompt SFT dataset for training an LLM to route prompts to the
right diffusion model at inference time. Given a concept and a target_model, the
model learns to emit the correct prompt dialect (FLUX T5-XXL prose, SDXL dual-clip
tokens, a compact caption, or steering modifiers).
Routing lives in the instruction prefix, not in a nested output object — keeping the
LoRA's task simple and maximizing structural diversity for generalization.… See the full description on the dataset page: https://huggingface.co/datasets/Limbicnation/dual-stream-image-prompts.ImageGenEditPromptCorpus
Image Gen & Edit Prompt Corpus
This is a dataset that contains several image generation (T2I) and image edition (I2I) datasets' prompt corpus.
This dataset is designed to align different text encoder (text condition) for image generation task.
Ingredients
carpedkm/RCEdit-500K, license CC BY-NC 4.0
jackyhate/text-to-image-2M, license MIT
OmniGen2/X2I2, license Apache 2.0
WINDop/OpenGPT-4o-Image, license Apache 2.0
code-cinema-image-animee
Code du cinéma et de l'image animée, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-cinema-image-animee.CuPer_Images
CuPer_Images
Resumen del Dataset
CuPer_Images es un dataset multimodal especializado en culturas precolombinas de Perú, desarrollado por NovaIA, el laboratorio de inteligencia artificial de Grupo Neura. Este dataset fue utilizado como parte del entrenamiento de Amaru, un modelo de lenguaje de propósito general con conocimientos profundos en las civilizaciones ancestrales del Perú.
Información del Dataset
Tamaño total: 4,247 muestras
División: 3,396 muestras… See the full description on the dataset page: https://huggingface.co/datasets/NovaIALATAM/CuPer_Images.clinical-narrative-image-integrity-v0.1Clinical Narrative–Image Integrity v0.1
Goal
Test whether narrative context forces the model to invent image findings
Test whether the model can keep image evidence and story context separate
Detect image-grounding failures that look fluent but are structurally wrong
What it measures
hallucinationResponse asserts a positive finding not supported by image facts
contradictionResponse asserts a finding that directly conflicts with a negative image fact
missed_requiredResponse fails to mention… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-narrative-image-integrity-v0.1.MM-MathInstruct-longest-20k-solutions-with-images
MM-MathInstruct Longest 20K Solutions with Images
This dataset contains the top 20,000 samples from MathLLMs/MM-MathInstruct selected by solution length, filtered to include only samples with valid images.
Dataset Structure
Usage
Source
This dataset is derived from MathLLMs/MM-MathInstruct by selecting the 20,000 samples with the longest solution text.
License
Apache 2.0 (inherited from source dataset)
Images-Diffusion-Prompt-Style
Image Diffusion Prompt Style
High-quality synthetic prompts for image diffusion models, optimized for Flux, Z Image, and Qwen.
Dataset Structure
Column
Type
Description
style_name
string
Short descriptive name
prompt_text
string
Full prompt with quality tokens
negative_prompt
string
Artifacts to avoid
tags
list
Lowercase keywords
compatible_models
list
Target models
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/Limbicnation/Images-Diffusion-Prompt-Style.bio-image-analysis-qa
Dataset Card for bio-image-analysis-qa
This dataset contains questions and answers for analysing biological microscopy imaging data using python.
Dataset Details
Dataset Description
Questions and answers provided in this repository are centered around the topic, how to process imaging data using Python.
Curated by: Robert Haase
License: CC-BY 4.0
Dataset Sources and Processing
This dataset was derived from the Bio-image Analysis Notebooks which… See the full description on the dataset page: https://huggingface.co/datasets/haesleinhuepf/bio-image-analysis-qa.finance-legal-mrc-with-images
🧾 finance-legal-mrc-with-images (tableqa/test)
Multimodal-ready table image dataset designed for TIG (Table Information Generation) inputTIG 입력 전용 테이블 이미지 데이터셋 (VLM 활용 가능)
📦 Dataset Overview | 데이터셋 개요
Feature
Description (EN)
설명 (KR)
🧩 Split
tableqa / test
tableqa / test 스플릿
📄 Total Rows
1,197 (unique tables only)
총 1,197건 (중복 제거된 고유 테이블 기준)
🖼️ Image Format
PNG (rendered from raw HTML tables)
원본 HTML 테이블을 PNG 이미지로 렌더링
🔗 Source
Derived from… See the full description on the dataset page: https://huggingface.co/datasets/didi0di/finance-legal-mrc-with-images.Images
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Johnnyboystar/Images.multilingual-image-annotations-text
Multilingual Image Annotations (Text Only)
Text-only companion to Reubencf/multilingual-image-annotations. Same rows, same google/gemma-4-31B-it annotations, but the image and boxed_image columns are removed so the dataset is small and loadable without binary image bytes.
Stats
Rows: 464
Detection-applicable: 273 (58%)
Languages: en, es, fr, hi, zh, ar, pt
Schema
Column
Type
Notes
image_id
string
UUID/stem of original file
description_en
string… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/multilingual-image-annotations-text.
