CoolFace
Modelpublic

dougvk/Unlimited-OCR-RDNA4

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes
Model Card

Unlimited-OCR RDNA 4

Single-GPU Baidu Unlimited-OCR inference for AMD RDNA 4 (gfx1200 / gfx1201) using native ROCm PyTorch, verified end to end on gfx1201.

This project does not redistribute or rename Baidu's weights. It provides a reproducible runtime around the original BF16 checkpoint: an isolated AMD wheel bootstrap, pinned model preparation, audited custom-code patching, single-GPU selection, local-only model resolution during inference, atomic output, bounded PDF handling, diagnostics, and validation evidence.

Status: beta. Verified on 31 July 2026 with an AMD Radeon RX 9070 XT (gfx1201). This is not an AMD or Baidu project.

Verified configuration

ComponentVerified value
GPUAMD Radeon RX 9070 XT, 16 GB
GPU allocationOne isolated GPU
OSUbuntu 24.04.4
Kernel6.17.0-40-generic
Host ROCm7.2.1
PyTorch2.9.1 + ROCm 7.2.1
Transformers4.57.1
Matplotlib3.10.8
Model revision07dea832e22aefee32ad281d4b80551282e1c168
Weight precisionOriginal BF16; no quantization
Peak allocated VRAM8.25 GiB on the 1600×2000 validation page

The initial validation recovered headings, reading order, a structured table, euro amounts, a formula, and an exact checksum. Repeated greedy generation was byte-identical after deterministic tag cleanup. See docs/VALIDATION.md for the evidence and its limits.

Quick start

The bootstrap downloads checksum-pinned AMD ROCm wheels into your XDG cache, installs every other runtime/build dependency from requirements/bootstrap.lock with hashes required, builds this checkout without build isolation, and creates a repository-local virtual environment. It does not use sudo, alter /opt/rocm, or touch the system Python.

bash
git clone https://huggingface.co/dougvk/Unlimited-OCR-RDNA4
cd Unlimited-OCR-RDNA4

./scripts/bootstrap-rocm.sh
.venv/bin/unlimited-ocr-rdna4 prepare
.venv/bin/unlimited-ocr-rdna4 run --input page.png

The model download is approximately 6.78 GB. The verified AMD wheel set is approximately 1.9 GB before installation.

PDFs are rendered and parsed page by page:

bash
.venv/bin/unlimited-ocr-rdna4 run \
  --input document.pdf \
  --output document.md \
  --max-pages 10 \
  --dpi 200

Commands

prepare

Downloads the pinned original checkpoint, fully hashes the 6.67 GB safetensors file and every behavior-defining model, configuration, tokenizer, and index file, preserves Baidu's original model source, and applies a narrowly scoped audited patch.

bash
unlimited-ocr-rdna4 prepare
unlimited-ocr-rdna4 prepare --dry-run --json
unlimited-ocr-rdna4 prepare --model-dir /data/models/unlimited-ocr

Preparation is serialized and idempotent. It uses a unique partial directory, validates the exact manifest and file set, and refuses unknown, extra, symlinked, or mismatched files instead of executing them.

doctor

Reports the active PyTorch/HIP stack, visible devices, RDNA 4 architecture, BF16 capability, and model readiness.

bash
unlimited-ocr-rdna4 doctor
unlimited-ocr-rdna4 doctor --device 0 --require-model
unlimited-ocr-rdna4 doctor --json

run

Parses one image or PDF and publishes untrusted model output atomically. The content may include Markdown and raw HTML; do not render it in a privileged origin without sanitization.

bash
unlimited-ocr-rdna4 run --input scan.png
unlimited-ocr-rdna4 run --input scan.png --output scan.md --device 0
unlimited-ocr-rdna4 run --input report.pdf --start-page 21 --max-pages 10

Important flags:

FlagDefaultMeaning
--device$UNLIMITED_OCR_DEVICE or 0ROCm ordinal or stable ROCr UUID
--modegundamgundam for detailed single-page parsing; base for lower-detail input
--max-length4096Total input + output sequence limit; maximum 32768
--dpi200PDF rendering resolution
--max-pages20Per-run PDF safety cap
--max-page-pixels60000000Per-page rendered-pixel cap
--max-total-pixels400000000Aggregate rendered-pixel cap
--max-page-rendered-mib512Per-page rendered temporary-byte cap
--max-rendered-mib2048Aggregate rendered temporary-byte cap
--forceoffReplace an existing output file
--jsonoffStable machine-readable summary on stdout
--quietoffSuppress progress diagnostics

OCR content is written to the output file. Progress and warnings go to stderr. Human or JSON summaries go to stdout.

GPU selection

PyTorch uses the CUDA-compatible API name on ROCm. --device sets ROCR_VISIBLE_DEVICES before PyTorch is imported, so the process sees one logical cuda:0 backed by the selected AMD card.

For ordinary one-GPU systems:

bash
unlimited-ocr-rdna4 run --input page.png --device 0

For multi-GPU systems, a stable ROCr UUID avoids dependence on enumeration order:

bash
unlimited-ocr-rdna4 run --input page.png --device GPU-0123456789abcdef

Run unlimited-ocr-rdna4 doctor before choosing a device. Do not assume another application's GPU numbering matches ROCm's ordinal numbering.

What the RDNA 4 adaptation changes

  • —Pins the Baidu checkpoint and verifies the complete expected local model tree before custom code loads.
  • —Uses AMD's production ROCm 7.2.1 PyTorch, torchvision, and Triton wheels for Python 3.12.
  • —Isolates one gfx1200/gfx1201 GPU before importing PyTorch.
  • —Replaces unsafe eval() calls in optional model-output geometry parsing with ast.literal_eval().
  • —Makes one internal mask transfer follow the active tensor device.
  • —Supplies the missing all-ones attention mask and pad token for single-sequence generation.
  • —Loads only the prepared local model during inference and enables Hugging Face offline mode.
  • —Avoids vLLM, SGLang, quantization, tensor parallelism, and unverified custom serving kernels.
  • —Removes only complete, exact layout sentinel pairs without rewriting recognized Unicode or TeX.
  • —Detects obvious terminal repetition and warns without silently rewriting recognition content.

The model weights and mathematical operators are unchanged. PyTorch's ROCm backend provides the RDNA 4 kernels.

Configuration

Precedence is flags, then environment variables, then XDG defaults.

Environment variablePurpose
UNLIMITED_OCR_DEVICEDefault ROCm ordinal or UUID
UNLIMITED_OCR_MODEL_DIRPrepared model directory
XDG_DATA_HOMEDefault model storage root
XDG_CACHE_HOMEWheel, Hugging Face, and temporary-work cache root
NO_COLORAccepted implicitly; the CLI currently emits no color

This project contains no analytics or telemetry code.

Safety and limitations

  • —The default 4096 sequence limit includes visual-prefill tokens. Dense pages may require a larger value.
  • —PDF mode uses permissively licensed PDFium bindings and renders one bounded page at a time. A failed run leaves no published partial document; retry with --start-page to resume manually.
  • —Generative OCR can omit or hallucinate content. Verify consequential documents against the source.
  • —Rotated text and repetitive pages are known upstream weak spots.
  • —Only the configuration above has completed the full repository acceptance test. Newer ROCm/PyTorch versions may work, but doctor reports them as unverified until measured.
  • —The hardware guard accepts gfx1200, but reports hardware_verified=false; only gfx1201 has completed this GPU gate.
  • —The bootstrap currently supports Linux x86_64, Python 3.12, and host ROCm 7.2.1.

See SECURITY.md before processing untrusted documents.

Development

CPU-only tests do not install PyTorch. The commands below are convenient for development; the release bootstrap is the hash-locked installation path.

bash
python3.12 -m venv .venv-dev
.venv-dev/bin/pip install -e '.[dev]'
.venv-dev/bin/ruff check .
.venv-dev/bin/ruff format --check .
.venv-dev/bin/pytest
.venv-dev/bin/python -m build
.venv-dev/bin/twine check dist/*

The GPU smoke gate is intentionally separate. It uses tracked, hash-checked fixtures; parses structured JSON; runs the image twice in fresh processes; and checks image/PDF output hashes, structure, stack identity, revision, VRAM, and process release. On the validated host, run it in a root-created private network namespace:

bash
VALIDATION_NETWORK_ISOLATED=1 \
UNLIMITED_OCR_DEVICE=GPU-0123456789abcdef \
AMD_SMI_GPU=3 \
./scripts/validate-smoke.sh

Credits and license

The runtime is MIT licensed. Baidu's original copyright and license are preserved; see THIRD_PARTY_NOTICES.md.