1. Why Standard Browser PDF Libraries Fail on Urdu and Arabic Script
Most lightweight JavaScript PDF libraries rely on built-in Type1 Standard Fonts (such as Helvetica or Times Roman) that only support the WinAnsi (Windows-1252 Latin-1) character set.
As soon as a user types or translates a single Urdu, Arabic, Persian, or Hindi character (such as "د" or "ع"), standard PDF encoders throw an immediate "WinAnsi cannot encode" exception, or render disconnected, reversed letters from left to right.
2. Our Dual-Pipeline RTL & Complex Script Solution
ScanifyAI inspects every text string, watermark, annotation block, and translated paragraph before generating the PDF:
- Automatic Unicode Detection: Whenever non-Latin characters (Code Point > 127) are detected in the title, body, watermark, or edit blocks, ScanifyAI automatically switches from WinAnsi fonts to our High-DPI Unicode Canvas Engine.
- Native Contextual Shaping & Ligatures: By leveraging the browser’s native HarfBuzz text shaping engine paired with Noto Naskh Arabic and Noto Nastaliq Urdu web fonts, connected Arabic/Urdu letters join naturally.
- Per-Line RTL Alignment: Each line is analyzed independently—Urdu and Arabic lines align right-to-left (RTL) at the right margin, while English numbers, email addresses, and reference codes align left-to-right (LTR).
Open the relevant ScanifyAI guide view or browse document and PDF tools.