datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
InkSight-DerenderingsPaper: arxiv.org/abs/2402.05804
Citation
If you find it useful for your research and applications, please cite using this BibTeX:
@article{
mitrevski2025inksight,
title={InkSight: Offline-to-Online Handwriting Conversion by Teaching Vision-Language Models to Read and Write},
author={Blagoj Mitrevski and Arina Rak and Julian Schnitzler and Chengkun Li and Andrii Maksai and Jesse Berent and Claudiu Cristian Musat},
journal={Transactions on Machine Learning Research}… See the full description on the dataset page: https://huggingface.co/datasets/Derendering/InkSight-Derenderings.vesuvius-ink-sweepsinkference-book-imagesSumi-e_no_kurozuappu_ink_painting_closeups
Sumi-e no Kurozuappu Ink Painting Closeups
Dataset Description
This dataset contains 1,250 synthetically generated images of Japanese Sumi-e ink paintings. The images were created to explore the capabilities of diffusion models in replicating the nuanced art style of traditional Japanese ink painting. The dataset is ideal for those interested in niche artistic model creation, generating LORAs, or simply enjoying and sharing these artistic expressions.
Example… See the full description on the dataset page: https://huggingface.co/datasets/takara-ai/Sumi-e_no_kurozuappu_ink_painting_closeups.ink3d-example-datavesuvius-eligible-meshes-ink9
ink_9um maps for the 340 eligible-scroll surface meshes
Correction, 2026-09-15
The direction labels here are swapped relative to the team's layer order.
These renders were made without --flip-normals, and the meshes write
outward normals like the team's, so the file named _reverse is the
forward pass in the team's convention and the unsuffixed file is the
reverse one. Both passes are present, so nothing is missing, but any
forward-vs-reverse comparison changes… See the full description on the dataset page: https://huggingface.co/datasets/pscamillo/vesuvius-eligible-meshes-ink9.inkmvgntinkslop-results
InkSlop Benchmark Results
Model evaluation results for the InkSlop Benchmark - a vibe-coded benchmark for spatial reasoning with digital ink.
Collection: InkSlop Benchmark
Contents
This dataset contains inference results and evaluation metrics for multiple VLMs across all InkSlop tasks:
overlap_easy / overlap_hard - Overlapped handwriting recognition
autocomplete_easy / autocomplete_hard - Handwriting autocompletion
derender_easy / derender_hard - Ink derendering (image… See the full description on the dataset page: https://huggingface.co/datasets/amaksay/inkslop-results.InkEraser
ynyg/InkEraser
筆跡擦除數據集,用於將含有筆跡的圖像(source)轉換為去除筆跡的目標圖像(target)。
文件結構
data/
train/
source/ # 原圖(含筆跡)
target/ # 目標(去筆跡)
val/
source/
target/
配對規則
source/ 與 target/ 中使用完全相同的文件名進行一一對應。
圖像格式:.jpg。
數據量(當前倉庫)
train:1511 對
val:115 對
合計:1626 對
使用示例(Python)
from huggingface_hub import snapshot_download
from pathlib import Path
from PIL import Image
root = Path(snapshot_download("ynyg/InkEraser", repo_type="dataset"))… See the full description on the dataset page: https://huggingface.co/datasets/ynyg/InkEraser.inkslop-mazes-hard
InkSlop Mazes Hard
Part of the InkSlop Benchmark a vibe-coded benchmark for spatial reasoning with digital ink.
Collection: InkSlop Benchmark
Task
Maze Solving: Given a maze image, find and draw the solution path as digital ink. This "hard" variant contains diverse synthetically generated maze families with varying visual styles.
Maze Families
bamboo - Bamboo forest style
cave_skeleton - Cave system layouts
floorplan_rooms - Architectural floor plans… See the full description on the dataset page: https://huggingface.co/datasets/amaksay/inkslop-mazes-hard.lunakr_ink_painting_05inkslop-autocomplete-hard
InkSlop Autocomplete Hard
Part of the InkSlop Benchmark a vibe-coded benchmark for spatial reasoning with digital ink.
Collection: InkSlop Benchmark
Task
Handwriting Autocompletion: Given a partial handwritten input, generate the completion as digital ink. This "hard" variant contains human-collected handwriting samples.
Data Format
This dataset contains two top-level directories:
original/ # Raw collected data
└── samples/
└──… See the full description on the dataset page: https://huggingface.co/datasets/amaksay/inkslop-autocomplete-hard.kr_ink_painting_07kr_ink_painting_08Angeli_inkTXkgbmFtZTogSSB0d2lzdCBhbmQgdHVybiwgeWV0IHN0YXkgdGhlIHNhbWUsIEEgc2h1ZmZsZWQg
c2VsZiwgYSB3b3JkeSBnYW1lLiBCcmVhayBtZSBhcGFydCwgSeKAmW0gc3RpbGwgaW4gcGxheSDi
gJREaWZmZXJlbnQgZm9ybXMsIGJ1dCBoZXJlIEkgc3RheS4
license: cc0-1.0
InkBench-deepseek-ocr
Document OCR using DeepSeek-OCR
This dataset contains markdown-formatted OCR results from images in NealCaren/InkBench using DeepSeek-OCR.
Processing Details
Source Dataset: NealCaren/InkBench
Model: deepseek-ai/DeepSeek-OCR
Number of Samples: 10
Processing Time: 1.9 min
Processing Date: 2025-10-22 21:01 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split: train
Batch Size: 1
Resolution Mode: base
Base Size: 1024
Image Size: 1024… See the full description on the dataset page: https://huggingface.co/datasets/NealCaren/InkBench-deepseek-ocr.ink_sketch_v1InkBenchkr_ink_wash_paintingkr_ink_painting_02Peterkr_ink_wash_painting_total_0927kr_ink_painting_04kr_ink_painting_03ink-style-imageskr_ink_painting_01twitter-Inkyung_slut-2024.11.14-1857067913095544889-4lIIHz1O1i6rv73g-part1ZPE-Ink-lane-state
ZPE-Ink Lane State Snapshot
Lane: zpe-ink
Portfolio: encoding
Canonical source repo: Zer0pa/ZPE-Ink (PUBLIC)
Snapshot date: 2026-05-12
Purpose: Durability snapshot of lane-agent receipts and the product-page draft for the ZPE-Ink lane, so the lane can be fully resumed from GitHub + Hugging Face after a local-disk wipe.
This dataset repository is not the codec, the codec proofs, or the public benchmark artifacts. Those all live in the GitHub repo Zer0pa/ZPE-Ink (PyPI: pip install… See the full description on the dataset page: https://huggingface.co/datasets/Zer0pa/ZPE-Ink-lane-state.kr_ink_painting_06
