us

98.83.72.38

Back
Blogs

Non-Latin Document Verification: How to Improve Accuracy Across Languages and Scripts

Non-Latin Document Verification: How to Improve Accuracy Across Languages and Scripts
Madiha Khatoon JUNE 3, 2026 13 minutes read

Non-Latin document verification requires script-aware OCR, document classification, Unicode normalization, transliteration, source reconciliation, and authenticity checks. Because coverage claims and aggregate accuracy scores conceal script-specific failures, businesses should test providers using representative documents, critical fields, onboarding conditions, and separate fraud controls before choosing a platform.

The Unicode Consortium released version 17.0 of the Unicode Standard on 9 September 2025. It added 4 scripts, bringing the total number of Unicode-encoded scripts to 172. However, that number is not a useful measure of document-verification coverage by itself. A provider does not need every script equally. It needs reliable performance in the countries, document types, and scripts your customers actually use.

Coverage claims can therefore hide the real problem. A system may list Arabic or Chinese as ‘supported’ while still failing to recognize a name, reversing a field, dropping a diacritic, or treating two legitimate spellings as if one were fraudulent. This guide explains how global IDs are read, where errors enter the process, and how to test a provider with real-world evidence that accurately reflects production conditions.

What is Non-Latin Document Verification, and Why Does It Affect Onboarding?

Non-Latin document verification is the process of reading and validating identity documents that contain scripts other than the Latin alphabet. Common examples include Arabic, Chinese Han characters, Cyrillic, Thai, Devanagari, Hangul, and Ge’ez. The same workflow may also need to handle Latin-based languages with complex diacritics, such as Vietnamese.

It helps to separate three tasks that are often mixed. OCR extracts visible text. Document verification checks whether the document and its data appear genuine and internally consistent. Identity verification then links the document to the person, often through a biometric or trusted data source. Meanwhile, AML screening is a downstream compliance check against sanctions, PEP, and other risk data. A good system connects these stages, but success at one stage does not prove success at the others.

This distinction matters commercially. If the system misreads a name or address, a genuine customer may be asked to retry or may be sent to manual review. However, if the matching rules are too loose, the business can also miss fraud or create noisy screening alerts. Reliable multilingual processing protects both conversion and decision quality. Focusing on accuracy and user experience helps ensure genuine users feel valued and understood.

How do Verification Systems Read Non-Latin Identity Documents?

A single universal OCR model does not usually recognize a global ID. Instead, a verification system coordinates several readers and checks. The exact path depends on whether the document contains only visible text or also includes a machine-readable zone, a barcode, or a contactless chip.

  1. Capture and classify the document. The system checks image quality, detects the document type and version, identifies field regions, and determines the writing direction. This prevents an Arabic or bilingual layout from being forced into a Latin, left-to-right template.
  2. Extract data from every available source. OCR reads the visual inspection zone. Meanwhile, a parser can read the MRZ or barcode, and an NFC reader can access a supported chip. These sources overlap, but they are not interchangeable. OCR reads printed fields, while chip authentication can provide cryptographic evidence about stored data.
  3. Normalize the values without erasing meaning. The system standardizes Unicode representation, reading direction, date formats, and digit styles. It should preserve the original script and diacritics, because aggressive conversion can turn two different names into the same string or change a legitimate name. Proper normalization is key to maintaining data integrity across scripts.
  4. Transliterate and reconcile name variants. When a name appears in both a national script and Latin letters, the system links the versions rather than requiring a character-for-character match. It also checks supporting attributes such as date of birth, document number, and nationality.
  5. Authenticate the document and make the decision. The workflow compares fields across sources, checks MRZ check digits when present, tests for document security and tampering signals, and validates the chip signature when NFC and trust data are available. It can then combine the result with face matching, database checks, or AML screening, as defined by the business’s risk policy.

Important: An MRZ is valuable, but it does not prove that a passport is genuine. Its check digits help detect reading errors and certain data inconsistencies. For an ePassport, passive authentication of the chip provides stronger evidence that the stated authority issued the stored data and that it has not been altered. ICAO’s Doc 9303 framework separates optical reading, chip data structures, and security mechanisms for this reason.

Which Scripts Create the Most Common ID-Reading Challenges?

Dozens of countries issue identity documents in non-Latin or bilingual formats. The table below is not a coverage list. Instead, it shows why each script needs its own extraction and matching tests.

Script Typical ID Examples Main Reading Challenge and Control
Arabic National IDs and passports in Saudi Arabia, the UAE, Qatar, and Egypt Arabic runs right-to-left, while embedded numbers run left-to-right. Letters also join and change form by position. Use bidirectional layout handling, document-specific field detection, and Arabic-trained extraction.
Chinese (Han) Mainland Chinese resident identity cards and passports A large character set and dense address fields make small recognition errors costly. Read the Chinese fields directly, then reconcile any Latin transcription instead of replacing the source text.
Cyrillic Russian, Bulgarian, and Serbian identity or travel documents Some Cyrillic and Latin letters look alike but have different Unicode values. Detect the script, normalize the text, and compare official transliterations as variants.
Thai Thailand’s national ID card and passport Vowels and tone marks can sit above or below consonants, and spaces do not always separate words. Preserve every mark and compare Thai and English name fields without assuming a character-for-character match.
Devanagari and other Indic scripts Aadhaar and Indian regional identity documents Conjunct characters and local-language fields vary by script and region. Test the exact languages used in the target states, not an aggregate ‘Indic’ score.
Hangul South Korean identity and travel documents Korean letters are grouped into syllable blocks, while romanization can vary. Retain the Hangul name and treat the Latin spelling as a linked representation.
Ge’ez (Ethiopic) Ethiopian identity documents and credentials The Ethiopic abugida is absent from many generic OCR datasets. Verify support for real Amharic or other Ethiopic-script fields, and for the exact document layouts accepted.
Vietnamese (extended Latin) Vietnam’s citizen identity card and passport Vietnamese is Latin-based, not non-Latin. However, stacked diacritics can be lost or decomposed differently in Unicode. Preserve the marks and normalize text before matching.

Arabic layouts also mix directions within the same line: the script runs right-to-left, while numbers normally run left-to-right. The W3C Arabic layout requirements describe this bidirectional behavior. Meanwhile, India’s UIDAI explains that Aadhaar data entered in English is transliterated into the selected local language, and users may need to correct the result. This is a useful reminder that bilingual fields are related representations, not guaranteed exact matches. UIDAI’s language and transliteration guidance makes that limitation explicit.

non-latin-document-verification

Where does Non-Latin Document Verification Fail?

Failures usually enter at five points: capture and classification, extraction, normalization and transliteration, cross-source validation, and downstream screening. A weak result early in the pipeline can therefore appear later as a false document mismatch or an AML alert.

Capture and classification: the wrong template creates the wrong crop

A document can fail before OCR reads a character. Glare, blur, a cropped edge, or the wrong document version template can move the expected field region. Right-to-left and mixed-direction layouts add another layer. Therefore, the system should classify the document and locate fields from the actual layout before applying script-specific recognition.

Extraction: OCR must preserve characters, marks, and field structure

Arabic joining, Thai tone marks, CJK character density, and Indic conjuncts create different recognition problems. A single aggregate language score cannot show which one is failing. Field boundaries matter as much as character accuracy. Reading the issuer’s address instead of the holder’s address still counts as a verification failure, even when every character is recognized correctly.

Normalization and transliteration: one person can have several valid spellings

Transliteration represents a name in another script; it does not translate the name’s meaning. According to ICAO Doc 9303, mandatory passport data written in a non-Latin script must be provided with a Latin transcription or transliteration. The MRZ uses a restricted character set built around A to Z, 0 to 9, and the filler symbol <. As a result, diacritics may be removed, names may be shortened to fit the field, and different authorities may use different permitted representations.

For example, Muhammad, Mohammed, and Mohamed can all refer to the same Arabic name. A raw exact-string comparison can therefore reject a genuine user. However, accepting every similar name is also unsafe. The system should compare the original script, the official Latin representation, and supporting identity fields together.

Cross-source validation: readable data is not the same as an authentic document

A forged image can still contain perfectly readable text. Therefore, OCR accuracy must be measured separately from document authenticity performance. Where a passport has an MRZ and a chip, the system should compare the visible data with the MRZ and chip data, and validate the chip’s digital signature when the required trust chain is available. For documents without a chip, the workflow needs other forensic and consistency checks.

Screening: poor input creates noisy matches

Screening should not rely on one transliterated string. It may need native-script names, Latin variants, aliases, and other identifiers. OFAC’s own Sanctions List Search scores names with string and phonetic algorithms, Jaro-Winkler and Soundex, and OFAC does not provide a recommended confidence rating because the right setting depends on the risk and facts of each search. OFAC’s sanctions-search FAQs show why thresholds must be tested rather than copied.

Better extraction reduces avoidable variation before screening begins. Meanwhile, contextual attributes such as date of birth, country, document number, and address help analysts distinguish a genuine match from a similar name. This lowers false-positive noise without weakening the review.

Practitioners report that gap directly. In Shufti’s Voice of Customer research across more than 600 organisations, 11% said they cannot reliably tell a true match from a false one, and common names were among the reasons given. A name transliterated into a restricted character set becomes a common name by construction, which is why the original script and the supporting identifiers have to survive all the way to the screening engine.

Five failure points in non-Latin document verification, from document capture to downstream screening.

How to Evaluate Multilingual OCR without Relying on Coverage Claims

A supported-language count is a discovery signal, not proof of production accuracy. A defensible pilot should use documents that reflect the actual customer mix and report results by country, document type, version, and script.

Historical rejects are useful because they expose hard cases. However, they are a biased sample and cannot predict overall performance on their own. Combine them with a representative, consented, and securely handled production sample, plus known fraud and image-quality edge cases.

  1. Ask for Per-Script and Per-Document Results: Do not accept one global OCR number. Require field-level accuracy for the scripts, countries, and document versions that matter to your rollout.
  2. Measure Critical Fields Separately: A missed accent and a wrong document number do not create the same risk. Track exact and normalized accuracy for names, dates of birth, document numbers, expiry dates, and addresses.
  3. Test the Full Journey, not OCR in Isolation: Measure first-attempt completion, false rejection, manual-review rate, and median and tail latency. Then trace each failure back to capture, extraction, matching, or authenticity checks.
  4. Challenge Transliteration and Reconciliation: Use genuine passports or IDs where the national script and Latin fields differ legitimately. Confirm that the system links the variants without hiding the original name.
  5. Separate OCR Accuracy from Fraud Detection: A clean extraction score says nothing about tampering detection; test manipulated documents, replayed images, and chip or MRZ inconsistencies as a separate control set.
  6. Inspect the Benchmark Method: Ask how ground truth was created, which fields were scored, whether diacritics counted, how low-confidence reads were handled, and whether human review was included in the reported accuracy and latency.

How Shufti Reads Global Identity Documents Across Scripts

Shufti built its document verification engine in-house to read multilingual and non-Latin identity documents.  Its OCR model head-to-head benchmarks against Google Vision. Shufti reads Arabic at 92.17% compared to Google’s 90.24%, Vietnamese at 96.79% compared to 82.36%, and Chinese, Japanese, and Korean scripts at 86.87% compared to 82.89%. The system handles native character sets such as Devanagari, Ge’ez, Arabic, and Kanji, and uses recursive OCR for non-Latin scripts.

Shufti has global coverage with local depth across 240+ countries and territories and 10,000+ actively processed document types for 150+ languages. These figures show scale; however, buyers should still request the benchmark method, run a pilot in each country, and document the routes they serve. That is the fairest way to connect published accuracy figures to your own document mix, image quality, and risk policy.

For businesses expanding across the Gulf, Asia, Eastern Europe, or Africa, the value is practical. As a glocal platform managing the full compliance lifecycle, from sign-up and onboarding through authentication, monitoring, and remediation, Shufti extracts local scripts, reconciles Latin representations, and passes structured data into document checks and screening in one workflow. Therefore, teams can reduce avoidable retries while keeping the original identity data available for audit and review.

See how Shufti reads your hardest-market documents on real files, then book a demo.

Frequently Asked Questions

Why does OCR struggle with Arabic identity documents?

Arabic is written right-to-left, while numbers are usually written left-to-right. Many letters also join and change form depending on their position within a word. Therefore, Arabic document verification needs bidirectional layout handling, Arabic-aware recognition, and document-specific field extraction. Generic left-to-right templates are not enough.

Can one verification system handle both Latin and non-Latin documents?

Yes, but that does not mean one OCR model should process every script. A global verification system can coordinate document classifiers, script-specific OCR, MRZ or barcode parsers, NFC readers, normalization rules, and matching logic. The buyer should test the combined system by script and document type rather than accepting a single engine-level claim.

Is Vietnamese a non-Latin script?

No. Vietnamese uses the Latin alphabet with additional diacritics. However, those marks carry meaning and may be represented in multiple Unicode forms. A system that drops or mishandles them can change the name it reads, so Vietnamese should still be tested as a distinct OCR and normalization case.

Does the passport MRZ remove the need for non-Latin OCR?

No. The MRZ provides standardized Latin-script data for machine reading, but it contains only selected fields and may shorten names. The visual zone still matters, and many national identity cards lack an MRZ. Moreover, MRZ check digits support consistency checks; they do not prove that the document was genuinely issued.

What is the best way to compare multilingual OCR providers?

Build a consented test set that reflects your real countries, document versions, scripts, and image quality. Add difficult historical failures and known fraud cases, then measure field accuracy, first-attempt completion, manual review, latency, and fraud detection separately. Finally, review the errors by script, because an aggregate score can hide the corridor that matters most.

Disclaimer: The information provided here is for general informational purposes only and should not be treated as legal, regulatory, or business advice. Shufti Pro Limited accepts no liability for decisions or actions taken in reliance on this information.

Join the
Shufti Sphere Newsletter

Get the latest trends, insights, and expert opinions on KYC, AML, fraud prevention, and more, straight to your inbox.

    Pitch a piece and get a verified byline in the Media room.

    Partnership Inquiries?
    Email us at [email protected]

    iBeta Level 1 — ISO 30107-3 Compliant iBeta Level 2 — ISO 30107-3 Compliant iBeta Level 3 — ISO 30107-3 Compliant PCI DSS SOC 2 Type 2 GDPR GDPR Fundamentals — Quality Guild ISO 27001:2022 KJM Age Verification CCPA / CPRA Cyber Essentials Cyber Essentials Plus
    Copyright © 2026 Shufti. All rights reserved.