datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
image-text-pairs-ja-cc0-2
はじめに
このデータセットは画像生成で日本語を生成したいときに使うデータセットです。
ライセンス
CC-0です。著作権を放棄して使いやすくしました。
作り方の概要
gpt-oss-20bを使って、約8万個からなる単語集兼短文集を作りました。
その文章をPillowとPythonでランダム要素を入れながら100万枚と10万枚でレンダリングしました。
フォントはNoto Sans JPなのでライセンス的には問題ないと思います。
image-text-pairs-ja-cc0
Japanese Glyph Images with English Captions (CC0)
This dataset contains Japanese glyph images rendered with black text on white background.
Each .png image has a corresponding .txt file with an English caption:
This image is saying "<Japanese>". The background is white. The letter is black.
Structure
train/ — PNG images and matching TXT captions (same base filename)
provenance/assets_registry.csv — Fonts and license info
LICENSE.txt — CC0-1.0 license
Generation… See the full description on the dataset page: https://huggingface.co/datasets/alfredplpl/image-text-pairs-ja-cc0.
