datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Argimi-Ardian-Finance-10k-text-image
The ArGiMI Ardian datasets : text and images
The ArGiMi project is committed to open-source principles and data sharing.
Thanks to our generous partners, we are releasing several valuable datasets to the public.
Dataset description
This dataset comprises 34,000 financial annual reports, written in English, meticulously
extracted from their original PDF format to provide a valuable resource for researchers and developers in financial
analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.image-text-pairs-ja-cc0-2
はじめに
このデータセットは画像生成で日本語を生成したいときに使うデータセットです。
ライセンス
CC-0です。著作権を放棄して使いやすくしました。
作り方の概要
gpt-oss-20bを使って、約8万個からなる単語集兼短文集を作りました。
その文章をPillowとPythonでランダム要素を入れながら100万枚と10万枚でレンダリングしました。
フォントはNoto Sans JPなのでライセンス的には問題ないと思います。
text-to-image-2M
text-to-image-2M: A High-Quality, Diverse Text-to-Image Training Dataset
Overview
text-to-image-2M is a curated text-image pair dataset designed for fine-tuning text-to-image models. The dataset consists of approximately 2 million samples, carefully selected and enhanced to meet the high demands of text-to-image model training. The motivation behind creating this dataset stems from the observation that datasets with over 1 million samples tend to produce better… See the full description on the dataset page: https://huggingface.co/datasets/rivisia/text-to-image-2M.image-text-pairs-ja-cc0
Japanese Glyph Images with English Captions (CC0)
This dataset contains Japanese glyph images rendered with black text on white background.
Each .png image has a corresponding .txt file with an English caption:
This image is saying "<Japanese>". The background is white. The letter is black.
Structure
train/ — PNG images and matching TXT captions (same base filename)
provenance/assets_registry.csv — Fonts and license info
LICENSE.txt — CC0-1.0 license
Generation… See the full description on the dataset page: https://huggingface.co/datasets/alfredplpl/image-text-pairs-ja-cc0.text-to-image-cleanedtext-to-image-2M
text-to-image-2M: A High-Quality, Diverse Text-to-Image Training Dataset
Overview
text-to-image-2M is a curated text-image pair dataset designed for fine-tuning text-to-image models. The dataset consists of approximately 2 million samples, carefully selected and enhanced to meet the high demands of text-to-image model training. The motivation behind creating this dataset stems from the observation that datasets with over 1 million samples tend to produce better… See the full description on the dataset page: https://huggingface.co/datasets/Junaid7188/text-to-image-2M.qwen-image-text-renderingtext-image-finetune-1k-v2Amazon_image_text_pair
