datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TextEdit
TextEdit: A High-Quality, Multi-Scenario Text Editing Benchmark for Generation Models
Danni Yang,
Sitao Chen,
Changyao Tian
If you find our work helpful, please give us a ⭐ or cite our paper. See the InternVL-U technical report appendix for more details.
🎉 News
[2026/03/06] TextEdit benchmark released.
[2026/03/06] Evaluation code and initial baselines released.
[2026/03/06] Leaderboard updated with latest models.
📖 Introduction… See the full description on the dataset page: https://huggingface.co/datasets/opencompass/TextEdit.LLaVA-Med-60K-IM-text
LLaVA-Med-60K-IM-text
This dataset is a text format of llava_med_instruct_60k_inline_mention.json.
We built this dataset using the Meta-Llama-3-70B-Instruct, and the instruction we used is: Rewrite the question-answer pairs into a paragraph format (Do not use the words 'question' and 'answer' in your responses):.
PMC articles that failed to download are excluded.
Non-medical images (e.g., diagrams) are excluded in an automatic way.
Despite these efforts, this dataset is not… See the full description on the dataset page: https://huggingface.co/datasets/myeongkyunkang/LLaVA-Med-60K-IM-text.Low-light_Scene_Text_Dataset
Low-light Scene Text Dataset
This repository provides a low-light scene text recognition dataset for studying text recognition under challenging illumination conditions. The dataset is designed to support research on Low-light Scene Text Recognition (LLSTR), where text images may suffer from low contrast, noise, uneven illumination, blur, and other degradations commonly observed in nighttime or poorly lit environments.
The dataset contains two main parts:
LSTR: a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/lumimusta/Low-light_Scene_Text_Dataset.korean-text-rendering-data
한글 텍스트 렌더링 학습 데이터
이미지 안에 정확한 한글 텍스트를 렌더링하는 능력 개선을 위해 만들어진 합성(synthetic) 이미지-프롬프트 데이터셋입니다. 2026년 5월~7월에 걸쳐 진행된 세 차례의 별도 학습 이터레이션에서 나온 데이터를 통합했습니다.
총 79,460장, 2개 config(콘텐츠 유형)로 구성. 각 config는 독립적으로 로드할 수 있습니다.
from datasets import load_dataset
ds = load_dataset("<repo_id>", name="diagram") # 유형별로 필요한 것만
이 릴리즈는 순수 한글 타이포그래피 학습에 초점을 맞춰 atomic_text(99.4% 한글)와
diagram(100% 한글) 두 유형만 포함합니다. 둘 다 코드·템플릿 기반 결정론적 생성이라
외부 생성형 서비스에 의존하지 않고, 라이선스 문제가 없습니다. "프롬프트 안 인용부호=정답
텍스트" 컨벤션은 둘 다… See the full description on the dataset page: https://huggingface.co/datasets/fasoo/korean-text-rendering-data.PMC-VQA-text
PMC-VQA-text
This dataset is a text format of PMC-VQA.
We built this dataset using the Meta-Llama-3-70B-Instruct, and the instruction we used is: Rewrite the question-answer pairs into a paragraph format (Do not use the words 'question' and 'answer' in your responses):.
train_text.json corresponds to the train.csv and train_2.csv splits in the PMC-VQA dataset.
Samples with two or more question-and-answer pairs were selected.
Citation
If you find this dataset useful… See the full description on the dataset page: https://huggingface.co/datasets/myeongkyunkang/PMC-VQA-text.Text2Receipt
Text2Receipt
Messy free-text Hebrew income notes -> valid, complete Israeli fiscal documents (receipts & tax invoices).
Live demo (Space): yonilev/Text2Receipt
Dataset: yonilev/Text2Receipt
Dataset Creation
A synthetic corpus from a deterministic, rule-based generator plus a bounded LLM-paraphrase layer, so the ground truth is exact by construction.
Pipeline
Scenario sampling - category, issuer status, document type, client type, year, payment… See the full description on the dataset page: https://huggingface.co/datasets/yonilev/Text2Receipt.stargate_s04e01_100topkdiverse_text2vid
trending-text-to-image
CivitAI Improved Prompts Dataset
This dataset contains trending AI-generated images from CivitAI with Flux-improved prompts for better generation results.
Dataset Format (JSONL)
Each line contains a JSON object with:
id: Original image ID from CivitAI
improved_prompt: Flux-enhanced version of the prompt
category: Automatically determined theme category
All original CivitAI metadata including:
Original prompt and negative prompt
Model information
Image URL and… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/trending-text-to-image.atomic2023-small_text2imagecomponent-image-texteis-text250
EIS-Text250: 1970s U.S. Environmental Impact Statements (text-only)
Per-page OCR/extraction text for 250 scanned 1970s U.S. federal
Environmental Impact Statements (EIS) from the Northwestern University
Library collection — the text-only companion to
Windsao/eis-subset50
(which carries full page images for a 50-doc subset). Built to test how current
models handle long, dense, historical government text: mean ~300 pages/doc,
1970s typewriter prose, OCR noise from degraded… See the full description on the dataset page: https://huggingface.co/datasets/Windsao/eis-text250.texturecan
Dataset Card for TextureCan Textures
Dataset Summary
This dataset contains 4,037 texture images from texturecan.com. It includes textures of various materials such as brick, paper, fabric, metal, wood, stone, and other surfaces. The original archives were downloaded, unpacked, and images were compressed using PNG optimization and JPEG quality compression (90%) to reduce file size while maintaining good quality.
Languages
The dataset is monolingual:
English… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/texturecan.testing-text-image2bitcoin-news-articles-text-corporacoco_text_traintest2017llava_finetuning_dataset_for_text_extraction
