CoolFace
Modelpublic

cstr/dbnet-ic15-GGUF

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
1likes237downloads
Model Card

DBNet ResNet-18 ICDAR 2015 — GGUF

Text detection model for CrispEmbed. Detects text regions (words/lines) in document images, scene photos, and screenshots.

Architecture: DBNet (Differentiable Binarization Network) with ResNet-18 backbone, FPNC neck, and probability map head. All BatchNorm folded into Conv at export time.

Source: MMOCR `dbnet_resnet18_fpnc_1200e_icdar2015` (Apache 2.0). This is the ICDAR 2015 Incidental Scene Text / Challenge 4 detector benchmark, not the separate ICDAR2015-TextSR dataset. Reported benchmark scores are P=0.885, R=0.758, H=0.817.

Dataset attribution: ICDAR 2015 Incidental Scene Text and the MMOCR dataset metadata. The similarly named ICDAR2015-TextSR dataset is a different dataset with separate ODbL terms and was not used as the training source for this detector.

Model Variants

VariantSizeCosine vs F32max_absDetection quality
F3247 MBbaselinereference
F1624 MB1.0000002.4e-3identical
Q8_013 MB1.0000001.9e-2identical (21/21 regions)
Q4_K7 MB1.0000002.9e-1identical (21/21 regions)

All variants detect the same text regions with scores within 0.005 of F32. Parity validated via per-pixel diff harness against PyTorch reference.

Recommended: Q80 (13 MB) or F16 (24 MB) for production fidelity. Q4K is a debug/experimental override only; its numerical parity does not change the recognition model's separate precision policy.

Usage

Pair with a TrOCR recognition model (cstr/trocr-small-printed-GGUF) for a complete OCR pipeline.

CLI

bash
# Full OCR pipeline (detect + recognize)
crispembed --det dbnet-ic15-q8_0.gguf \
    -m trocr-small-printed-q8_0.gguf \
    --ocr document.png

# JSON output
crispembed --det dbnet-ic15-q8_0.gguf \
    -m trocr-small-printed-q8_0.gguf \
    --ocr document.png --json

C API

c
#include "crispembed.h"

void *ctx = crispembed_ocr_init("dbnet-ic15-q8_0.gguf",
                                 "trocr-small-printed-q8_0.gguf", 4);
int n;
const crispembed_ocr_result *r = crispembed_ocr(ctx, "image.png", &n);
for (int i = 0; i < n; i++)
    printf("(%g,%g): %s\n", r[i].x, r[i].y, r[i].text);
crispembed_ocr_free(ctx);

Architecture

Input image (resized, padded to 32x)
  |
  +-> ResNet-18 backbone (stem + 4 stages x 2 BasicBlocks)
  |     Stage 0: 64ch, stride 4   Stage 1: 128ch, stride 8
  |     Stage 2: 256ch, stride 16  Stage 3: 512ch, stride 32
  |
  +-> FPNC neck (FPN-Cat variant)
  |     4x lateral 1x1 conv -> top-down upsample+add
  |     4x smooth 3x3 conv (256->64) -> concat 4x64=256ch
  |
  +-> DBHead probability branch
        3x3 conv (256->64) + ReLU
        ConvTranspose2d (64->64, k=2, s=2) + ReLU
        ConvTranspose2d (64->1, k=2, s=2) + sigmoid

Post-processing: binarize at 0.3 -> connected components -> bbox extraction with unclip expansion (ratio 1.5). Output sorted in reading order.

12.2M parameters. All BatchNorm pre-folded into Conv/ConvTranspose weights.

Conversion

bash
pip install gguf numpy torch mmengine

# Download MMOCR checkpoint
wget -q "https://download.openmmlab.com/mmocr/textdet/dbnet/dbnet_resnet18_fpnc_1200e_icdar2015/dbnet_resnet18_fpnc_1200e_icdar2015_20220825_221614-7c0e94f2.pth"

# Convert
python models/convert-dbnet-to-gguf.py \
    --checkpoint dbnet_resnet18_fpnc_1200e_icdar2015_20220825_221614-7c0e94f2.pth \
    --output dbnet-ic15-f32.gguf

# Quantize
crispembed-quantize dbnet-ic15-f32.gguf dbnet-ic15-q8_0.gguf q8_0
crispembed-quantize dbnet-ic15-f32.gguf dbnet-ic15-q4_k.gguf q4_k

License and attribution

The converted model artifacts are released under Apache-2.0, with source attribution to MMOCR. Retain the dataset attribution above when redistributing this model or derived artifacts.

Provenance and EU AI Act Art. 53 note

  • Upstream model: open-mmlab/mmocr.
  • Upstream licence: apache-2.0. This repository redistributes under the same terms; it grants no rights the upstream licence does not.
  • What was done here: format conversion and/or quantisation only (GGUF). No training, no fine-tuning, no merging, no distillation, no change to architecture, vocabulary or capability. Only the numeric representation of the upstream weights differs.
  • Training data: documented — where it is documented at all — by the upstream provider; see the upstream model card. No training data was used, added or selected by this repository.
  • Provider status: under Regulation (EU) 2024/1689 the upstream authors remain the provider of this model. Converting the serialisation format does not make this repository the provider of a new general-purpose AI model, and no such claim is made. Questions about training content, copyright policy or model capability belong upstream.