CoolFace
Modelpublic

tori29umai/rtdetrv4-x-manga109s_v2

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
7likes
Model Card

RT-DETRv4-X Manga109-s v2

[日本語](#日本語) | [English](#english)

漫画ページから コマ枠 (frame) / 人物 (body) / 台詞 (text) / 顔 (face) の 4 クラスを検出する RT-DETRv4 (X size) モデル。Manga109-s の見開きを単ページに分割した 15,263 枚 (単ページ 14,796 + 見開き保持 467) で 30 epoch ファインチューンしたもので、v1 (3 クラス) に `face` クラスを追加したバージョンです。


日本語

概要

漫画ページから 4 クラス (コマ枠・人物・台詞・顔) を検出する RT-DETRv4 (X size) モデル:

  • —0: body — 人物 (キャラクター本体)
  • —1: text — 台詞 / 吹き出し領域
  • —2: frame — コマ枠
  • —3: face — 顔 (v2 で追加)

Manga109-s の見開きを単ページ分割した 15,263 枚 (単ページ 14,796 + 見開き保持 467) で 30 epoch ファインチューン、DINOv2 ViT-B/14 特徴蒸留付き。bbox がセンター線をまたぐページは分割せず見開きのまま保持しているため、学習・推論ともに単ページと見開き両方に対応します。v1 では body のみで人物全体を取っていましたが、v2 では body (体) と face (顔) を分離検出することで、表情解析やキャラクタークロップなど顔単位の下流処理に対応できるようにしています。ComfyUI ワークフロー、自動化されたコマ単位処理パイプライン、漫画ドメインの研究を想定しています。

ベースアーキテクチャRT-DETRv4 X-size (HGNetv2-B5 backbone + DFINETransformer decoder)
蒸留教師モデルDINOv2 ViT-B/14 (Apache 2.0)
学習データManga109-s — 商用利用許諾済 87 タイトル
入力解像度1280 × 1280
クラス数4 (body / text / frame / face)

検出例

bbox の色分け: <span style="color:#9acd32">黄緑 = frame (コマ)</span> / <span style="color:#1e90ff">青 = body (人物)</span> / <span style="color:#dc143c">赤 = text (セリフ)</span> / <span style="color:#ff69b4">ピンク = face (顔)</span>

[image]

完成原稿 (ペン入れ済み) の例。コマ・人物・台詞・顔の 4 クラスすべて高精度で取れています。

[image]

ラフな手書きネーム (下描き / ストーリーボード) の例。学習データは完成原稿だけですが、線画の途中段階でもコマ枠・人物・台詞・顔をある程度検出できます。

精度

Manga109-s validation split (1,212 ページ, 30,252 bbox) で評価 (best_stg2, epoch 27):

指標値
平均 mAP75.0%
平均 AP5096.0%
平均 AP7579.3%
平均 AR10081.6%
AP (small)21.4%
AP (medium)56.2%
AP (large)80.6%

検出ヒット率を表す AP50 が 96% に達しており、実用上ほぼ取りこぼしなし。残りの伸びしろは IoU 厳格化 (AP75) と、特に小サイズ bbox (細かい台詞や群衆中の小さな顔) であり、検出漏れではなく bbox 枠の精度向上が今後の改善ポイント。

ファイル

ファイル説明
model.onnxONNX (opset 17, 静的入力 1×3×1280×1280)

推論

ONNX グラフの入出力:

  • —入力: images (float32, NCHW, [0, 1] に正規化), orig_target_sizes (int64, [N, 2] = [width, height])
  • —出力: labels (int, [N, 300]), boxes (float32, [N, 300, 4], 元画像座標の xyxy), scores (float32, [N, 300])

onnxruntime での最低限のサンプル:

python
import numpy as np
import onnxruntime as ort
from PIL import Image, ImageDraw

CLASS_NAMES = {0: "body", 1: "text", 2: "frame", 3: "face"}
INPUT_SIZE = 1280
CONF_THRESHOLD = 0.5

session = ort.InferenceSession(
    "model.onnx",
    providers=["CUDAExecutionProvider", "CPUExecutionProvider"],
)

image = Image.open("page.jpg").convert("RGB")
W, H = image.size

# 前処理: 1280x1280 にリサイズ → CHW float32 [0, 1]
resized = image.resize((INPUT_SIZE, INPUT_SIZE), Image.BILINEAR)
arr = np.asarray(resized, dtype=np.float32) / 255.0
arr = arr.transpose(2, 0, 1)[None]  # 1x3x1280x1280
orig_size = np.array([[W, H]], dtype=np.int64)

labels, boxes, scores = session.run(
    None, {"images": arr, "orig_target_sizes": orig_size}
)
labels, boxes, scores = labels[0], boxes[0], scores[0]

# 信頼度閾値で絞る (boxes は既に元画像の座標で出てくる)
keep = scores >= CONF_THRESHOLD
print(f"Detected {int(keep.sum())} objects")
for cid, (x1, y1, x2, y2), s in zip(labels[keep], boxes[keep], scores[keep]):
    print(f"  {CLASS_NAMES[int(cid)]:5s}  conf={s:.3f}  "
          f"bbox=({x1:.0f},{y1:.0f},{x2:.0f},{y2:.0f})")

# 可視化
draw = ImageDraw.Draw(image)
colors = {0: "blue", 1: "red", 2: "yellow", 3: "hotpink"}
for cid, (x1, y1, x2, y2) in zip(labels[keep], boxes[keep]):
    draw.rectangle([x1, y1, x2, y2], outline=colors[int(cid)], width=3)
image.save("output.png")

CPU で動かす場合は providers=["CPUExecutionProvider"] のみで OK。GPU 用には onnxruntime-gpu をインストール。

学習設定

エポック30 (flat 15 + cosine 11 + no_aug 4)
バッチサイズ16 (single GPU)
オプティマイザAdamW — lr=2.5e-4, backbone lr=2.5e-6, weight_decay=1.25e-4
拡張Mosaic / RandomPhotometricDistort / RandomZoomOut / RandomIoUCrop, Mixup (epoch 2–15)
蒸留DINOv2 ViT-B/14 特徴蒸留, loss_distill weight 20 (adaptive)
train / valtrain: 15,263 枚 (単ページ 14,796 + 見開き保持 467) / val: 1,212 枚 (単ページ 1,116 + 見開き保持 96)、bbox は 370,738 / 30,252
クラス別 bbox 数 (train)body 109,480 / text 105,139 / frame 75,581 / face 80,538
クラス別 bbox 数 (val)body 8,645 / text 8,433 / frame 6,541 / face 6,633

Manga109-s の見開きページは原則センター線で単ページに分割して学習。ただし bbox がセンター線をまたぐページ (見開きを横断するコマやキャラクターを含むページ) は分割せず見開きのまま保持してアノテーションを残す mixed-mode。これは推論時の典型的なシナリオ (1 度に 1 ページ) と入出力を揃えつつ、学習中に「センター線をまたぐ正解 bbox」を欠落させないため。

v1 との差分

項目v1v2
クラス数3 (body / text / frame)4 (body / text / frame / face)
train bbox 総数290,200370,738 (+ face 80,538)
val bbox 総数23,61930,252 (+ face 6,633)
想定ユースケースコマ・人物・台詞のレイアウト解析上記 + 顔単位のクロップ・表情解析

body の定義は v1 と同一 (人物全体の bbox)。v2 では face を追加で別 bbox として持っており、body 内に face がネストする形になります (face ⊂ body ではなく独立したアノテーションとして検出される)。

ライセンス・帰属表示

モデル本体: Apache License 2.0

本モデルは [Manga109-s](http://www.manga109.org/ja/download_s.html) を学習データとして使用しています。Manga109-s の規約に従って以下を明示します:

  • —データセット本体は同梱しません。Manga109-s の取得は公式サイトからの正規入手に従ってください。
  • —このモデルを使って Manga109-s 収録漫画画像の 複製・改変を商材化することは規約により禁止されています。
  • —Manga109-s に基づく結果を公表する際は下記 2 論文の引用が必要です。

引用 (BibTeX)

bibtex
@article{multimedia_aizawa_2020,
    author={Kiyoharu Aizawa and Azuma Fujimoto and Atsushi Otsubo and Toru Ogawa and Yusuke Matsui and Koki Tsubota and Hikaru Ikuta},
    title={Building a Manga Dataset ``Manga109'' with Annotations for Multimedia Applications},
    journal={IEEE MultiMedia},
    volume={27},
    number={2},
    pages={8--18},
    doi={10.1109/mmul.2020.2987895},
    year={2020}
}

@article{mtap_matsui_2017,
    author={Yusuke Matsui and Kota Ito and Yuji Aramaki and Azuma Fujimoto and Toru Ogawa and Toshihiko Yamasaki and Kiyoharu Aizawa},
    title={Sketch-based Manga Retrieval using Manga109 Dataset},
    journal={Multimedia Tools and Applications},
    volume={76},
    number={20},
    pages={21811--21838},
    doi={10.1007/s11042-016-4020-z},
    year={2017}
}

謝辞

  • —RT-DETRv4 / D-FINE (Apache 2.0) — 本モデルのベースアーキテクチャと学習コード
  • —DINOv2 (Apache 2.0) — 蒸留教師モデル
  • —HGNetv2 (Apache 2.0) — backbone

English

Overview

RT-DETRv4 (X-size) finetuned on Manga109-s for 4-class object detection on Japanese manga pages:

  • —0: body — characters / human figures
  • —1: text — dialogue balloons & text regions
  • —2: frame — panel borders
  • —3: face — character faces (added in v2)

This is the v2 release of RT-DETRv4-X Manga109-s, with an additional `face` class on top of the original 3 classes. While v1 captured whole characters via body, v2 separates body (full figure) and face (head only), enabling face-level downstream tasks such as expression analysis or character cropping. Trained on 15,263 images (14,796 single pages + 467 retained spreads, split from Manga109-s) for 30 epochs with DINOv2 ViT-B/14 feature distillation. Pages whose bboxes cross the centerline are kept as full spreads, so the model handles both single pages and spreads at inference time. Intended for ComfyUI workflows, automated panel-level pipelines, and manga-domain research.

Base architectureRT-DETRv4 X-size (HGNetv2-B5 backbone + DFINETransformer decoder)
Distillation teacherDINOv2 ViT-B/14 (Apache 2.0)
Training dataManga109-s — 87 commercially-licensed titles
Input resolution1280 × 1280
Number of classes4 (body / text / frame / face)

Examples

bbox color coding: <span style="color:#9acd32">yellow-green = frame (panel)</span> / <span style="color:#1e90ff">blue = body (character)</span> / <span style="color:#dc143c">red = text (dialogue)</span> / <span style="color:#ff69b4">pink = face</span>

[image]

A finished, inked manga page. All four classes (panel / character / dialogue / face) are picked up with high precision.

[image]

A rough hand-drawn "name" (storyboard / pre-inking sketch). Although the training data only contains finished manga, the model still recognises panels, characters, dialogue regions and faces reasonably well at the rough-draft stage.

Performance

Evaluated on the Manga109-s validation split (1,212 pages, 30,252 boxes) with best_stg2 (epoch 27):

MetricValue
mean mAP75.0%
mean AP5096.0%
mean AP7579.3%
mean AR10081.6%
AP (small)21.4%
AP (medium)56.2%
AP (large)80.6%

AP50 reaches 96% — virtually no missed detections in practical use. The remaining headroom is in IoU strictness (AP75) and especially in small-size bboxes (tiny dialogue balloons, faces in crowd scenes). The bottleneck is bbox-tightness, not recall.

Files

FileDescription
model.onnxONNX, opset 17, static 1×3×1280×1280 input

Inference

The ONNX graph exposes:

  • —inputs: images (float32, NCHW, normalised to [0, 1]), orig_target_sizes (int64, [N, 2] = [width, height])
  • —outputs: labels (int, [N, 300]), boxes (float32, [N, 300, 4], xyxy in original image coordinates), scores (float32, [N, 300])

Minimum working example with onnxruntime:

python
import numpy as np
import onnxruntime as ort
from PIL import Image, ImageDraw

CLASS_NAMES = {0: "body", 1: "text", 2: "frame", 3: "face"}
INPUT_SIZE = 1280
CONF_THRESHOLD = 0.5

session = ort.InferenceSession(
    "model.onnx",
    providers=["CUDAExecutionProvider", "CPUExecutionProvider"],
)

image = Image.open("page.jpg").convert("RGB")
W, H = image.size

# Preprocess: resize to 1280x1280, CHW float32 in [0, 1]
resized = image.resize((INPUT_SIZE, INPUT_SIZE), Image.BILINEAR)
arr = np.asarray(resized, dtype=np.float32) / 255.0
arr = arr.transpose(2, 0, 1)[None]  # 1x3x1280x1280
orig_size = np.array([[W, H]], dtype=np.int64)

labels, boxes, scores = session.run(
    None, {"images": arr, "orig_target_sizes": orig_size}
)
labels, boxes, scores = labels[0], boxes[0], scores[0]

# Filter by confidence (boxes are already in original image coordinates)
keep = scores >= CONF_THRESHOLD
print(f"Detected {int(keep.sum())} objects")
for cid, (x1, y1, x2, y2), s in zip(labels[keep], boxes[keep], scores[keep]):
    print(f"  {CLASS_NAMES[int(cid)]:5s}  conf={s:.3f}  "
          f"bbox=({x1:.0f},{y1:.0f},{x2:.0f},{y2:.0f})")

# Visualise
draw = ImageDraw.Draw(image)
colors = {0: "blue", 1: "red", 2: "yellow", 3: "hotpink"}
for cid, (x1, y1, x2, y2) in zip(labels[keep], boxes[keep]):
    draw.rectangle([x1, y1, x2, y2], outline=colors[int(cid)], width=3)
image.save("output.png")

For CPU-only inference, use providers=["CPUExecutionProvider"]. For GPU, install onnxruntime-gpu.

Training

Epochs30 (flat 15 + cosine 11 + no-aug 4)
Batch size16 (single GPU)
OptimiserAdamW — lr=2.5e-4, backbone lr=2.5e-6, weight_decay=1.25e-4
AugmentationMosaic / RandomPhotometricDistort / RandomZoomOut / RandomIoUCrop, Mixup (epoch 2–15)
DistillationDINOv2 ViT-B/14 feature distillation, loss_distill weight 20 (adaptive)
Train / val splittrain: 15,263 images (14,796 single pages + 467 retained spreads) / val: 1,212 images (1,116 single pages + 96 retained spreads); 370,738 / 30,252 bboxes
Per-class bbox count (train)body 109,480 / text 105,139 / frame 75,581 / face 80,538
Per-class bbox count (val)body 8,645 / text 8,433 / frame 6,541 / face 6,633

Manga109-s spreads were split at the centerline into single pages prior to training, except when an annotated bbox crossed the centerline — in that case the page was retained as a spread (mixed mode). This keeps the inference contract single-page-friendly while preserving cross-spread groundtruth boxes (e.g. panels or characters that span both pages) instead of clipping them away.

Differences from v1

Itemv1v2
Number of classes3 (body / text / frame)4 (body / text / frame / face)
Train bbox total290,200370,738 (+ face 80,538)
Val bbox total23,61930,252 (+ face 6,633)
Use caseLayout analysis of panels / characters / dialogueAbove + face-level cropping & expression analysis

The definition of body is unchanged from v1 (full-figure bbox). v2 additionally emits face as a separate, independent bbox — face is not strictly nested under body; both are detected in parallel.

License & Attribution

Model: Apache License 2.0.

This model was trained on [Manga109-s](http://www.manga109.org/en/download_s.html), whose terms of use require the following acknowledgements:

  • —The dataset itself is not bundled with this release. Obtain Manga109-s through the official channel.
  • —Using this model to commercially redistribute or sell reproductions / derivatives of Manga109-s manga images is prohibited by the dataset terms.
  • —The two papers below must be cited when reporting results that depend on Manga109-s.

Citation

bibtex
@article{multimedia_aizawa_2020,
    author={Kiyoharu Aizawa and Azuma Fujimoto and Atsushi Otsubo and Toru Ogawa and Yusuke Matsui and Koki Tsubota and Hikaru Ikuta},
    title={Building a Manga Dataset ``Manga109'' with Annotations for Multimedia Applications},
    journal={IEEE MultiMedia},
    volume={27},
    number={2},
    pages={8--18},
    doi={10.1109/mmul.2020.2987895},
    year={2020}
}

@article{mtap_matsui_2017,
    author={Yusuke Matsui and Kota Ito and Yuji Aramaki and Azuma Fujimoto and Toru Ogawa and Toshihiko Yamasaki and Kiyoharu Aizawa},
    title={Sketch-based Manga Retrieval using Manga109 Dataset},
    journal={Multimedia Tools and Applications},
    volume={76},
    number={20},
    pages={21811--21838},
    doi={10.1007/s11042-016-4020-z},
    year={2017}
}

Acknowledgements

  • —RT-DETRv4 / D-FINE (Apache 2.0) — base architecture and training code
  • —DINOv2 (Apache 2.0) — distillation teacher
  • —HGNetv2 (Apache 2.0) — backbone