Images and scanned PDFs
PNG, JPEG and PDF up to 20 MiB and 20 pages, rendered at 72–300 DPI.
Razin OCR
Model
In development
Razin OCR reads text from images and scanned PDFs and runs on your own hardware. It's general-purpose; our focus is making it especially good at Arabic, faithful down to the dots and diacritics. The engine and its evaluation harness work today on pretrained models. Our own fine-tuned model comes next, trained on permissioned, hand-checked documents.
Quarterly report
Total due 1,250.75 AED
PNG, JPEG and PDF up to 20 MiB and 20 pages, rendered at 72–300 DPI.
UTF-8 text and JSON with page geometry, reading order, engine confidence, warnings and pinned model provenance.
Models are fetched once with checked hashes. Inference never sends a document to a hosted OCR API.
No string reversal, no folding of alef or hamza forms, diacritics kept. The metrics are strict about the same things.
Our first run, on 6 September 2026, measured Arabic: three pretrained configurations on the same 14 images, on CPU with two threads. Lower error is better.
| Configuration | Strict CER | Strict WER | Median s/image | Peak RSS MiB |
|---|---|---|---|---|
| Tesseract fast, ara+eng | 52.80% | 58.82% | 0.199 | 109.5 |
| Tesseract best, ara+eng | 42.03% | 48.04% | 0.288 | 148.1 |
| PaddleOCR PP-OCRv5 + Arabic recognizer | 15.23% | 23.53% | 1.888 | 571.4 |
14 synthetic diagnostic pages, 1,392 reference characters: correlated diagnostics, not 14 independent documents, and not evidence of accuracy on your documents. No training or fine-tuning has run.
| Case | Measured | What happened |
|---|---|---|
| Mixed Arabic, English and digits | 39.14% | Paddle's character error rate. An order number and an amount vanished; an email address came back corrupted. |
| Diacritics and letter distinctions | 15.00% | Paddle again. Vowel marks and the dots that separate letters are exactly what strict metrics count. |
| Lines Tesseract skipped | 2 of 4 | On a widely spaced printed page, both Tesseract configurations dropped the middle two lines entirely. |
UTF-8 text and JSON with page geometry, reading order, engine confidence, warnings and pinned model provenance.
{
"schema_version": 1,
"model": { "engine": "paddle", "version": "3.7.0", "device": "cpu" },
"pages": [{
"width": 1600, "height": 700,
"coordinate_space": "processed_page_pixels_top_left",
"regions": [{
"text": "Quarterly report 2026",
"polygon": [[113, 120], [880, 113], [881, 209], [114, 215]],
"confidence": 0.9999,
"granularity": "line",
"reading_order": 0
}, …],
"warnings": ["Confidence is engine-specific and uncalibrated."]
}]
}Collect 100 permissioned, manually checked pages from one document category, split by document.
Annotate regions and text, and score recognition on identical crops alongside whole pages.
Use development errors to choose between segmentation fixes, mixed-script recognition and fine-tuning.
Publish the model with results on a test set it never saw, next to these baselines.
We're looking for design partners with permissioned documents: invoices, forms, letters, records. You get early access and numbers on your own pages.