datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Anime-LineArt-Dataset
Anime Lineart Sketch Dataset
Dataset Summary
This dataset contains high-quality lineart sketch images automatically extracted from raw anime images using the LineartAnimeDetector model from ControlNet Annotators (lllyasviel/Annotators).
It is designed to support research and development in:
Anime-style sketch generation
Text-to-sketch pipelines
ControlNet conditioning
Sketch-to-image and image-to-sketch translation
The raw source images are sourced from the Anime Images… See the full description on the dataset page: https://huggingface.co/datasets/ityizNola/Anime-LineArt-Dataset.LineArtImageNet100
LineArtImageNet100 v1
Summary
LineArtImageNet100 v1 is a synthetic, validation-only line-art image dataset
paired to the ImageNet100 validation split.
Dataset scope
Split: validation only
Size: 5,000 images
Classes: 100 ImageNet classes
Images per class: 50
Pairing: exact source-keyed pairing to the ImageNet100 validation inventory
Labels and structure
Labels are inherited from the paired ImageNet100 validation inventory and are… See the full description on the dataset page: https://huggingface.co/datasets/ebykAI/LineArtImageNet100.sumerian-lineart-padded-resized-augcomics_dataset_lineart_1024danbooru-lineartHi3D_Lineartsumerian-lineart-padded-and-resizedcoloringzap-lineart-parameters
Photo-to-line-art parameter sweep (coloringzap)
Review date: 2026-09-15 — every number here was measured by
scripts/export-dataset.mjs, not estimated.
Which settings turn a photo into a clean, printable line drawing? This dataset
is the measured answer: 128 rows covering 32 parameter combinations across 4
synthetic image classes, plus the 8 test images that went in and came out.
These measurements were produced by the same open-source pipeline that runs in the browser at… See the full description on the dataset page: https://huggingface.co/datasets/w46881/coloringzap-lineart-parameters.line_art_drawing_prompts
Dataset Card for "line_art_drawing_prompts"
More Information needed
linear-ttt-dclm-100m-600m
Qwen-LaCT:完整 Stage 1 100M + Stage 2 600M 训练数据
本目录包含实际冻结数据池的完整 token 数据,可复现训练输入。100M / 600M 是阶段预算的简称。
文件
完整输入 tokens
字节数
stage1-100m-99975168tokens.int32.bin
99,975,168
399,900,672
stage2-600m-599916544tokens.int32.bin
599,916,544
2,399,666,176
数据来源:mlfoundations/dclm-baseline-1.0,revision a3b142c183aebe5af344955ae20836eb34dcf69b。
格式:小端有符号 int32,tokenizer 为原生 Qwen2.5-3B-Instruct。文件是原始 master 的精确字节切片。
使用
将目录中的全部文件下载到同一目录。
执行 sha256sum -c… See the full description on the dataset page: https://huggingface.co/datasets/guo925658/linear-ttt-dclm-100m-600m.sumerian-lineart-processed-for-trocrlineartthree_line_summarization_for_japanese_news_articlesライブドアニュースコーパスの3行要約データセットです。
Llama v2向けのプロンプトを追加して成形してあります。
学習に利用する際は、 [R_START] [R_END] をspecial tokenとして追加することを推奨します。
Number of rows: 3,907
Datasetは以下のリポジトリを利用してscrapeしました。
git@github.com:KodairaTomonori/ThreeLineSummaryDataset.git
comics_orig_lineartLinear_tok1Linear_tok2sportwear_lineart_dataset
Sportwears Mates Lineart Dataset
Now this is a different type of image dataset.
Here it's not beautiful images, nor taggings.
These are spreadsheets in multiple views of sportsman in sportswear.
Oh, I admit this soundtrack sounds epic now (non-commercial, not part of license, no derivative work allowed)
Title: Stand As One (Stadium Chant Edition)
Despite not being exactly good for image generative models (which generate images), the idea is the other end of such chain.
It… See the full description on the dataset page: https://huggingface.co/datasets/eastenddan/sportwear_lineart_dataset.linear-testLinear_tok0VanGogh_Gladiolas1886_vs_TreeOil_LinearTorqueAxis_18Tech_GraphStudy
Van Gogh Gladiolas (1886) vs The Tree Oil Painting
Linear Torque Axis Analysis & 18 Supreme Techniques Forensic Study
Dataset by Haruthai Muangbunsri🔬 Powered by AI Sunny: TorqueBrush Forensics | Project Evergreen🗓️ Published 2025 on Hugging Face
📌 Abstract
This dataset provides a high-resolution comparative forensic analysis between Vincent van Gogh’s Gladiolas (1886) and the Tree Oil Painting using a proprietary AI model based on 18 Supreme Techniques and… See the full description on the dataset page: https://huggingface.co/datasets/HaruthaiAi/VanGogh_Gladiolas1886_vs_TreeOil_LinearTorqueAxis_18Tech_GraphStudy.Linear_tok3This is a modified version of the training dataset from the BabyLM challenge (https://arxiv.org/pdf/2301.11796), which the original authors provided under an MIT license.
one-line-artdanbooru-lineartThis is a dataset of paired lineart and colorized images, useful for training lineart extraction or lineart colorization models. Both lineart and colorized images are actual posts on Danbooru and not synthesized.
There are 609 pairs of images. They were collected in the following way:
Find all posts with tag 'lineart', and without tag 'sketch', and has a parent, and the parent does not have tag 'lineart'.
Select the pairs where the parent and the child have almost the same aspect ratio.… See the full description on the dataset page: https://huggingface.co/datasets/woctordho/danbooru-lineart.lineartHi3D_LineartSemanticDiffusion-for-Line-Art-datasetnva-lineartline-artlineart_controlnetdanbooru-lineart-3k
