Skip to content

Razin OCR

Model

In development

Razin OCR, measured in the open.

Razin OCR reads text from images and scanned PDFs and runs on your own hardware. It's general-purpose; our focus is making it especially good at Arabic, faithful down to the dots and diacritics. The engine and its evaluation harness work today on pretrained models. Our own fine-tuned model comes next, trained on permissioned, hand-checked documents.

Quarterly report

Total due 1,250.75 AED

Faithful text, with its geometry.

Images and scanned PDFs

PNG, JPEG and PDF up to 20 MiB and 20 pages, rendered at 72–300 DPI.

Text plus structure

UTF-8 text and JSON with page geometry, reading order, engine confidence, warnings and pinned model provenance.

Local by design

Models are fetched once with checked hashes. Inference never sends a document to a hosted OCR API.

Arabic kept intact

No string reversal, no folding of alef or hamza forms, diacritics kept. The metrics are strict about the same things.

Where pretrained models stand.

Our first run, on 6 September 2026, measured Arabic: three pretrained configurations on the same 14 images, on CPU with two threads. Lower error is better.

ConfigurationStrict CERStrict WERMedian s/imagePeak RSS MiB
Tesseract fast, ara+eng52.80%58.82%0.199109.5
Tesseract best, ara+eng42.03%48.04%0.288148.1
PaddleOCR PP-OCRv5 + Arabic recognizer15.23%23.53%1.888571.4

14 synthetic diagnostic pages, 1,392 reference characters: correlated diagnostics, not 14 independent documents, and not evidence of accuracy on your documents. No training or fine-tuning has run.

The hard parts are the small marks.

CaseMeasuredWhat happened
Mixed Arabic, English and digits39.14%Paddle's character error rate. An order number and an amount vanished; an email address came back corrupted.
Diacritics and letter distinctions15.00%Paddle again. Vowel marks and the dots that separate letters are exactly what strict metrics count.
Lines Tesseract skipped2 of 4On a widely spaced printed page, both Tesseract configurations dropped the middle two lines entirely.

result.json

UTF-8 text and JSON with page geometry, reading order, engine confidence, warnings and pinned model provenance.

{
  "schema_version": 1,
  "model": { "engine": "paddle", "version": "3.7.0", "device": "cpu" },
  "pages": [{
    "width": 1600, "height": 700,
    "coordinate_space": "processed_page_pixels_top_left",
    "regions": [{
      "text": "Quarterly report 2026",
      "polygon": [[113, 120], [880, 113], [881, 209], [114, 215]],
      "confidence": 0.9999,
      "granularity": "line",
      "reading_order": 0
    }, …],
    "warnings": ["Confidence is engine-specific and uncalibrated."]
  }]
}

From baseline to a Razin model.

  1. 01

    Real pages

    Collect 100 permissioned, manually checked pages from one document category, split by document.

  2. 02

    Region labels

    Annotate regions and text, and score recognition on identical crops alongside whole pages.

  3. 03

    Targeted training

    Use development errors to choose between segmentation fixes, mixed-script recognition and fine-tuning.

  4. 04

    Held-out results

    Publish the model with results on a test set it never saw, next to these baselines.

Have Arabic documents that need reading?

We're looking for design partners with permissioned documents: invoices, forms, letters, records. You get early access and numbers on your own pages.