datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
62k-images-khmer-printed-dataset
62k Khmer-English Printed Dataset
This repository contains a dataset of Khmer and English printed text images for training, validation, and testing. The dataset is stored in parquet format and managed using Git Large File Storage (LFS).
Installation
Prerequisites
Before cloning this repository, make sure you have Git LFS installed:
Install Git LFS
Linux/macOS:curl -s https://packagecloud.io/install/repositories/github/git-lfs/script.deb.sh | sudo… See the full description on the dataset page: https://huggingface.co/datasets/SoyVitou/62k-images-khmer-printed-dataset.bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bghira/pseudo-camera-10k with responses/captions generated with gemini-2.0-flash-thinking-exp-1219.
The format should be similar to that of liuhaotian/LLaVA-Instruct-150K.
Images can be found in the images.zip folder. The zip also contains .txt captions for ease of use in non-VQA tasks.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Images/bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.BLUEX_without_images
BLUEX: A benchmark based on Brazilian Leading Universities Entrance eXams
BLUEX is a multimodal dataset consisting of the two leading university entrance exams conducted in Brazil: Convest (Unicamp) and Fuvest (USP), spanning from 2018 to 2024. The benchmark comprises 724 questions that do not have accompanying images. This task evaluates models on the multiple-choice questions from these exams. The model is given a question in Portuguese together with its answer choices and… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/BLUEX_without_images.recipe-synthetic-images-10k
Recipe PDF Dataset
A multimodal dataset of 10K+ recipes rendered as PDF images with full metadata.
Dataset Description
Each sample contains:
image: Recipe rendered as a styled PDF page (PNG, ~1654x2339px)
name: Recipe title
description: Recipe description
ingredients: List of ingredients
steps: Cooking instructions
nutrition: Nutritional values (calories, fat%, sugar%, sodium%, protein%, sat.fat%, carbs%)
random_reviews: User reviews
minutes: Cooking time
tags: Recipe… See the full description on the dataset page: https://huggingface.co/datasets/TurkishCodeMan/recipe-synthetic-images-10k.Handpicked-Images-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
Handpicked-Images-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
Some random images with responses/captions generated with gemini-2.0-flash-thinking-exp-1219.
The format should be similar to that of liuhaotian/LLaVA-Instruct-150K.
Images can be found in the images.zip folder. The zip also contains .txt captions for ease of use in non-VQA tasks.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Images/Handpicked-Images-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.frontend-react-dataset-no-images
Frontend React Dataset — No Images
Image-free derivative of Reubencf/frontend-react-dataset.
This version preserves all 1,000 examples while completely removing the screenshot column and its embedded image bytes.
Schema
Column
Type
Description
description
string
Description or prompt for the interface
response
string
React/TSX implementation response
from datasets import load_dataset
ds = load_dataset("Reubencf/frontend-react-dataset-no-images")… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/frontend-react-dataset-no-images.MM-MathInstruct-longest-20k-solutions-with-images
MM-MathInstruct Longest 20K Solutions with Images
This dataset contains the top 20,000 samples from MathLLMs/MM-MathInstruct selected by solution length, filtered to include only samples with valid images.
Dataset Structure
Usage
Source
This dataset is derived from MathLLMs/MM-MathInstruct by selecting the 20,000 samples with the longest solution text.
License
Apache 2.0 (inherited from source dataset)
CuPer_Images
CuPer_Images
Resumen del Dataset
CuPer_Images es un dataset multimodal especializado en culturas precolombinas de Perú, desarrollado por NovaIA, el laboratorio de inteligencia artificial de Grupo Neura. Este dataset fue utilizado como parte del entrenamiento de Amaru, un modelo de lenguaje de propósito general con conocimientos profundos en las civilizaciones ancestrales del Perú.
Información del Dataset
Tamaño total: 4,247 muestras
División: 3,396 muestras… See the full description on the dataset page: https://huggingface.co/datasets/NovaIALATAM/CuPer_Images.Images-Diffusion-Prompt-Style
Image Diffusion Prompt Style
High-quality synthetic prompts for image diffusion models, optimized for Flux, Z Image, and Qwen.
Dataset Structure
Column
Type
Description
style_name
string
Short descriptive name
prompt_text
string
Full prompt with quality tokens
negative_prompt
string
Artifacts to avoid
tags
list
Lowercase keywords
compatible_models
list
Target models
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/Limbicnation/Images-Diffusion-Prompt-Style.finance-legal-mrc-with-images
🧾 finance-legal-mrc-with-images (tableqa/test)
Multimodal-ready table image dataset designed for TIG (Table Information Generation) inputTIG 입력 전용 테이블 이미지 데이터셋 (VLM 활용 가능)
📦 Dataset Overview | 데이터셋 개요
Feature
Description (EN)
설명 (KR)
🧩 Split
tableqa / test
tableqa / test 스플릿
📄 Total Rows
1,197 (unique tables only)
총 1,197건 (중복 제거된 고유 테이블 기준)
🖼️ Image Format
PNG (rendered from raw HTML tables)
원본 HTML 테이블을 PNG 이미지로 렌더링
🔗 Source
Derived from… See the full description on the dataset page: https://huggingface.co/datasets/didi0di/finance-legal-mrc-with-images.Images-Diffusion-Prompt-Style
Image Diffusion Prompt Style
High-quality synthetic prompts for image diffusion models, optimized for Flux, Z Image, and Qwen.
Dataset Structure
Column
Type
Description
style_name
string
Short descriptive name
prompt_text
string
Full prompt with quality tokens
negative_prompt
string
Artifacts to avoid
tags
list
Lowercase keywords
compatible_models
list
Target models
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/Pratofeitoo/Images-Diffusion-Prompt-Style.
