Arabic handwritten VLM

Hi everyone,

Based on your experience, which trained VLM performs best for extracting text from handwritten Arabic forms (e.g., applications, questionnaires, administrative documents)?

I’m mainly looking for the highest possible accuracy in:

  • Arabic handwritten text recognition
  • Structured form understanding (fields, tables, checkboxes)
  • Mixed printed and handwritten Arabic content
  • Ideally offline/self-hosted inference

Have you tested any models such as Qwen-VL, Florence, PaliGemma, or Arabic-specific OCR/VLM models on this type of document? If possible, I’d appreciate benchmark results, model recommendations, and fine-tuning advice.

Thank you!

For now, I tried a few things:


If the main goal is the highest practical accuracy on handwritten Arabic forms, I would not start by looking for one VLM that does everything. I would first separate page/field localization from handwriting recognition, and compare that modular route against an all-in-one document VLM on the same forms.

My default would currently be:

Mostly fixed templates?
├─ Yes
│  └─ register page → known field ROIs → specialist recognizers
│
└─ No / variable layouts
   └─ detect text/line regions → crop
      ├─ handwriting → Arabic HTR model
      ├─ printed text → normal OCR
      ├─ tables → table branch
      └─ checkboxes/radio → control branch

For the handwriting stage, Arabic-English-handwritten-OCR-v3 is the strongest practical starting point I found/tested so far. If you already have clean, pre-segmented single-line crops, AMIDDA Line4B is also very interesting: it is explicitly designed for line-level Arabic-script HTR rather than page OCR.

For a one-model/document-parser baseline, I would also test PaddleOCR-VL-1.6. It is a compact 0.9B document VLM, supports Arabic among 109 languages, and is designed for text, tables, formulas, charts, handwriting, etc. I would treat it as the all-in-one control, rather than assuming in advance that the modular pipeline must win.

I also put together an executed T4 example of the variable-layout branch:

Executed Colab notebook: docTR → Sherif v3, Transformers v5

It runs without providing any data: it builds a small demo form from public KHATT line samples, detects line regions with docTR 1.1, sends the crops to Sherif v3, and exports JSON, CSV, crops, and an annotated page. The executed version uses Transformers 5.17 on a T4.

For fine-tuning, I would first identify which stage is actually failing:

wrong crop / missed field       -> localization or registration
right crop, wrong transcription -> HTR
right text, wrong field         -> key/value association
table only fails                -> table structure
checkbox only fails             -> control detection

Then fine-tune the component that owns that error, rather than immediately fine-tuning a full-page VLM.

What I actually tested

Handwriting recognition

I ran a small same-harness directional comparison on Arabic handwriting crops. On that test, the mean CER was approximately:

Model Mean CER
Sherif v3 0.181
AMIDDA 0.282
PaddleOCR-VL-1.6 base 0.535

I would not treat those numbers as a general leaderboard. It was a small local comparison intended to answer a narrower question: whether there was enough signal to prefer a specialist recognizer over a general page model for the handwriting branch.

There was.

Fixed-form routing

I also made a controlled form-routing probe using real KHATT handwriting lines inserted into synthetic forms, with clean/mild/strong geometric distortion.

For six pages, simple feature-based registration + homography recovered the canonical form with:

  • 6/6 successful registrations
  • mean ROI IoU ≈ 0.998
  • mean ROI corner error ≈ 0.33 px

Again, this is an architecture diagnostic, not a benchmark on real administrative forms.

What it suggests is simply: if your forms are actually one or a few stable templates, there may be little benefit in making a VLM repeatedly infer where “Name”, “Address”, etc. are. Registering the page once and cropping known ROIs can remove an entire class of errors.

Variable-layout routing

On the same small controlled set:

  • docTR found a matching text region in all target handwriting regions; mean horizontal coverage was ≈ 0.842
  • PP-StructureV3 also hit all target regions; mean horizontal coverage was ≈ 0.954

I would not rank PP-StructureV3 above docTR from those figures:

  • only six synthetic pages were used;
  • this measures geometry/routing, not transcription;
  • the execution environments were different;
  • the goal was only to see whether both approaches could produce usable crops for an external HTR model.

Both did.

Model notes: Sherif, AMIDDA, Warraq, and generic VLMs

Sherif v3

Arabic-English-handwritten-OCR-v3 is based on Qwen2.5-VL-3B and is fine-tuned specifically for Arabic/English handwriting.

One reason I like it as a starting point here is that it is a full checkpoint, not an adapter that requires reproducing a particular PEFT attachment state. The published repository contains the full safetensors checkpoint.

I would treat the headline CER numbers in its model card as useful reference values rather than directly comparable benchmark numbers unless the datasets/protocols match your own evaluation.

The executed notebook above is also a useful sanity check that the current full checkpoint can be loaded under Transformers v5 and actually generate sensible output, rather than merely passing a loader check.

AMIDDA Line4B

AMIDDA Line4B is conceptually a very clean fit if your pipeline can already produce line crops.

Its current model card describes it as:

  • Qwen3.5-VL-4B based
  • trained on 53,420 lines from 8 corpora
  • LoRA-trained and merged back into the base weights
  • intended for pre-segmented single-line HTR
  • not a page-level OCR system
  • Transformers >= 5.9

That separation is useful: a layout/segmentation component can own geometry, while AMIDDA owns only the transcription problem.

Its published AMIDDA test-set result is 23.1% overall CER, with large differences across source corpora; for example the KHATT subset is reported at 9.8%. That variation is another reason I would test on your actual forms rather than select a model from one aggregate number.

Warraq / Fanar adapters

Warraq’s public evaluation report is also worth reading.

Using one common 1,038-line KHATT protocol, that report gives roughly:

  • Warraq multi-source: 3.86% CER
  • Warraq KHATT-specific: 4.11%
  • Sherif v3: 5.86%
  • generic Qwen2.5-VL-7B: 50.85%
  • Fanar base zero-shot: 50.98%

The important caveat is in the report itself: the Warraq variants are trained on KHATT (or KHATT plus other handwriting), whereas the other community models in that table are evaluated zero-shot on KHATT.

So I would not read that table as “Warraq is universally better than Sherif”. I would read it as strong evidence for a more useful point:

task-specific Arabic HTR adaptation can matter enormously compared with using a generic VLM as-is.

That is directly relevant if you are considering Qwen-VL, Florence, PaliGemma, etc. The question is not only which backbone is strongest; the handwriting-specific adaptation and target distribution can dominate the result.

Generic VLMs

I would still keep generic Qwen/other VLMs in the evaluation set if convenient, especially because they may be good at semantic structure.

But for the actual cursive transcription branch, I found much stronger reasons to start from an Arabic HTR fine-tune than from an untouched general-purpose VLM.

I did not find sufficiently comparable evidence to make Florence or a generic PaliGemma variant my first recommendation for this particular task.

Page parsing: docTR, PP-StructureV3, and PaddleOCR-VL

docTR

The runnable example uses docTR only as the geometry/front-end stage.

As of docTR 1.1, the project has expanded beyond basic OCR with layout analysis, table structure recognition, reading-order-aware exports, and a CLI.

For this notebook I deliberately use less than that: I only need line geometry, then I discard the general recognizer as the final answer and give each crop to the specialist HTR model.

That makes the boundary easy to understand and easy to replace.

PP-StructureV3

PP-StructureV3 is a richer alternative when you need document structure rather than just text lines.

It is useful for things such as:

  • layout regions
  • OCR polygons/coordinates
  • tables
  • reading order
  • structured JSON/Markdown output

That makes it attractive as a router/orchestrator even if you do not want to use its built-in recognizer for handwritten fields.

I did get it working in a controlled probe, but the Colab runtime was more version-sensitive than docTR. In particular, PaddlePaddle 3.3.1 CPU inference hit a oneDNN/PIR runtime issue that was avoided with:

FLAGS_enable_pir_api=0
enable_mkldnn=False

There is a related PaddleOCR issue here: PaddleOCR #18119.

That is why I used docTR in the small public notebook: it made the example simpler and more reproducible, not because I think PP-StructureV3 should be ruled out.

PaddleOCR-VL-1.6

PaddleOCR-VL-1.6 is the branch I would use as the all-in-one baseline.

The current PaddleOCR release describes it as a 0.9B document-parsing VLM with support for 109 languages, including Arabic, and handling of text, tables, formulas, charts, handwriting, and historical documents.

Its full document pipeline also combines the VLM with layout analysis, so this is much closer to the “one model/system for the entire form” idea than a line HTR model.

Also, one correction that matters for fine-tuning: current PaddleOCR-VL pipeline documentation does provide a supervised fine-tuning path for the VLM through ERNIEKit. The limitation currently documented is that fine-tuning of the layout analysis and ranking models is not supported.

So if PaddleOCR-VL is already close on your forms, fine-tuning it is a real option; I would just compare that cost against fine-tuning only the specialist recognizer in a modular pipeline.

Why I would not optimize only for CER

For forms, transcription CER is only one failure surface.

A useful reference here is KITAB-Bench, an Arabic OCR/document-understanding benchmark with 8,809 samples across 9 domains and 36 subdomains, including handwritten material, structured tables, and other document elements.

What I find useful about KITAB-Bench here is less “which model won” and more the decomposition of the problem.

For a form pipeline I would track at least:

localization / crop accuracy
        +
handwriting transcription
        +
field association
        +
table structure
        +
checkbox/control state
        +
final structured-record accuracy

A model can have a good CER and still produce a bad form record.

For example:

  • correct value, wrong field → OCR is fine, association failed
  • correct table text, wrong cell → recognition is fine, structure failed
  • empty field hallucinated as text → blank handling failed
  • checkbox interpreted as a character → wrong component owns the problem

Tables

I would treat table structure and text recognition as separate contracts.

Recognizing all strings on a table is not enough if row/column/cell assignment is wrong.

Checkboxes/radio buttons

If these are important, I would route them separately and return something explicit such as:

checked
unchecked
uncertain

rather than expecting the HTR model to turn every visual element into a text token.

Blank fields

I would also distinguish:

blank

from:

recognition_failed

Those two states may look similar in a final JSON record, but operationally they mean very different things.

How I would evaluate and fine-tune it

I would first collect a small representative target set, rather than trying to infer the answer from unrelated leaderboards.

Ideally include some combination of:

  • clean scans
  • phone photos
  • perspective/skew
  • weak illumination
  • dense handwriting
  • sparse handwriting
  • empty fields
  • Arabic + Latin + digits
  • tables
  • checkboxes/radio buttons

Then run the same pages through at least:

all-in-one document VLM
vs
modular parser/detector + specialist HTR

and score each stage separately.

For handwriting:

  • CER
  • WER if useful
  • exact field-value match

For the complete form:

  • region recall / crop coverage
  • correct key/value association
  • table/control accuracy
  • exact final field value
  • exact structured-record accuracy

Then the fine-tuning decision becomes much easier.

If crops are bad, training the recognizer will not fix the real problem.

If crops are good but transcription is poor, that is where HTR fine-tuning is justified.

If transcription is correct but fields are mismatched, the next work belongs in association/KIE, not OCR.

If only one form type has a problem, you may not need to touch the shared recognizer at all.

For PaddleOCR-VL specifically, the current docs point to ERNIEKit for VLM SFT. For Qwen-family HTR models there are also ordinary LoRA/QLoRA training routes, but I would only introduce that complexity once the target-page evaluation says the HTR stage is the actual bottleneck.

Arabic-specific correctness: RTL is not simply 'reverse the string'

One easy source of hidden errors is mixed-direction text.

Arabic fields commonly contain:

  • Arabic words
  • Latin abbreviations
  • phone numbers
  • dates
  • IDs
  • model/product codes

The current Unicode Bidirectional Algorithm (UAX #9, Revision 52) exists precisely because strings containing RTL and LTR content have a logical storage order and a separate visual display order.

So I would not “fix RTL” by blindly reversing OCR strings.

That can make an output look better on screen while silently corrupting identifiers, numbers, or embedded Latin text.

For evaluation, I would normalize Unicode deliberately and compare logical strings; leave visual bidi rendering to the display layer.

Transformers v5 and adapter compatibility

The public notebook uses a full Sherif checkpoint under Transformers 5.17.

I added more checks than just “from_pretrained() returned without throwing”:

  • the repository revision is checked for full model weights rather than adapter-only files
  • missing keys are reported
  • unexpected keys are reported
  • loader errors are reported
  • the current Qwen2.5-VL model.language_model.layers hierarchy is checked
  • a real weight tensor is checked for finite/nonzero values
  • generation is actually run

The executed notebook completed with:

missing keys:    0
unexpected keys: 0
loader errors:   0

and produced actual transcriptions.

That matters because adapter checkpoints deserve a stricter test across major library versions.

I ran into a case where a PEFT/LoRA adapter could appear to load and appear active, yet the trained tensors were not attached to the intended current module paths after a Transformers hierarchy change.

So for a major-version migration, I would treat:

"adapter loaded successfully"

as only the first check.

For an adapter I would additionally verify:

checkpoint keys
    ↓
current module mapping
    ↓
adapter tensors actually materialized
    ↓
output changes with adapter enabled vs disabled

This is mostly a future-reader note if you stay with a full merged checkpoint such as Sherif v3 or the current merged AMIDDA weights, but it is worth remembering if you experiment with LoRA repositories later.

For reference, the current Transformers multimodal flow is documented under the multimodal chat-template API.

So, if I had to reduce all of this to three starting points:

fixed/semi-fixed forms
    -> registration + known ROIs + specialist HTR

variable layouts
    -> detector/parser + specialist HTR
       (the runnable docTR -> Sherif example is one version)

one-model baseline
    -> PaddleOCR-VL-1.6

Then I would compare those routes on a few representative target pages at both transcription level and final field/record level.

That should tell you much more than choosing a model from a generic OCR leaderboard alone.

Thank you John for the very detailed explanation, i have started with the same recommanded model which looks good, will start following the steps you have described above, thank you again