Arabic Slugify
Slugs and transliteration for Arabic and Urdu text.
npm install @devix-labs/arabic-slugify
Twenty-one million downloads a week go to slug libraries that get Arabic script wrong. Two of them split Persian and Urdu words in half, because a zero-width joiner holds one word together and a kashida is a stretch rather than a space. All of them discard the vowel marks that would make a romanisation readable. And all of them insist on romanising at all, which for Arabic produces a string that means nothing to anyone. This keeps the script when you want a readable URL, uses the vowel marks when an author has written them, unifies spellings towards the content's own language, and adds the collision handling romanised Arabic badly needs. 2.0 kB, zero dependencies.
What you get
It can keep the script
Every browser has handled Unicode paths for years. /مرحبا-بالعالم is a better URL for Arabic content than mrhba-balaalm, and no other library offers it.
It does not split words in half
A zero-width non-joiner is how Persian writes inside a word. Two libraries with 8 million downloads a week between them turn میروم into my-rwm — two words. So does a kashida.
It uses the vowel marks
When an author has vocalised the text, مَرْحَبًا can be marhaban and مُحَمَّد can be muhammad. The most used alternative discards them and returns mrhba either way.
Unification follows the language
Persian and Urdu keep ی and ک; Arabic keeps ي and ك. Both spellings still collapse to one slug, so an article never gets two URLs — but a Persian URL is not spelled in Arabic letters.
uniqueSlug, which matters more here
Romanised Arabic collides far more than Latin does, because the short vowels that would tell two words apart are not written. Nothing else in this space handles it.
Direction marks never reach the URL
A right-to-left override can make a URL read as something other than what it is. It is stripped, along with every other invisible control character.
Three languages, properly
Positional readings for Persian's و and ی, a final heh as eh, and the letters پ چ ژ گ ٹ ڈ ڑ ں that an Arabic-only table drops entirely.
Honest about the limit
Unvocalised Arabic cannot be romanised into readable words — no library can do it without a dictionary. The docs say so, and point at keeping the script instead of pretending otherwise.
Arabic Slugify — overview
Why it exists
Twenty-one million downloads a week go to slug libraries, and every one of them gets Arabic script wrong in the same three ways.
They split words in half: a zero-width non-joiner is how Persian writes
inside a word — میروم is one word — and two libraries with eight million
downloads a week between them return my-rwm. The same happens on a kashida,
which is a decorative elongation rather than a space.
They throw away the information that would help: when an author has written
the vowel marks, مُحَمَّد can be muhammad; the most used alternative discards
them and returns mhmd regardless.
And they all insist on romanising, which for Arabic produces a string that means nothing to anybody:
| slugify | sindresorhus | speakingurl | transliteration | Laravel |
|---|---|---|---|---|
mrhba-balaalm |
mrhba-balealm |
mrhba-balaalm |
mrhb-blaalm |
mrhba-balaaalm |
What it does differently
script: 'keep'. Every browser has handled Unicode paths for years./مرحبا-بالعالمis a better URL for Arabic content than any romanisation, and no other library offers it.- Joiners and kashidas hold words together instead of breaking them.
- Vowel marks are used when they are there — including the detail that a shadda doubles the consonant, which after Unicode normalisation is ordered before it.
- Unification follows the language. Persian and Urdu keep ی and ک; Arabic keeps ي and ك. Both still collapse the other spelling, so one article never gets two URLs.
uniqueSlug()andprefersScript(), neither of which exists elsewhere.- Direction marks never reach the URL, because a right-to-left override can make one read as something it is not.
- 2.0 kB, no dependencies.
Not in 1.0
Vowel restoration. Turning مرحبا into marhaba needs a dictionary or
morphological analysis, and a table of guesses would be wrong often enough to be
worse than the consonants. The honest answer is to keep the script, and the docs
say so rather than hiding the limit.
Sun-letter assimilation. الشمس is romanised ash-shams by scholarly
convention and al-shms here. Predictable beats correct-for-some-words when the
output is a URL that has to stay stable.
How it compares
Questions
Why is مرحبا not marhaba?
Because Arabic script does not write short vowels. مرحبا is the consonants m-r-h-b plus a long alef; marhaba is not recoverable from it without a dictionary. Every library returns something like mrhba — slugify, speakingurl, transliteration and Laravel all give slightly different versions of the same unreadable string. That is why script: 'keep' exists.
Are Unicode URLs really safe to use?
Yes. Browsers have handled them for years, they are percent-encoded on the wire, and search engines index them. An Arabic slug is readable to the people the content is for, which is the entire purpose of a slug; a romanisation of it is readable to nobody.
Why does it change ی to ي — or not?
Unifying letter spellings is what stops one article having two URLs when it is typed twice. Which way it unifies follows the language: Arabic content unifies towards ي and ك, Persian and Urdu towards ی and ک. No other library makes the distinction, so a Persian slug comes out spelled in letters Persian does not use.
What is wrong with the zero-width joiner?
Nothing — it is doing its job. It holds one Persian or Urdu word together. The problem is libraries treating it as a separator: میروم is one word and they return my-rwm. The same mistake turns a kashida, which is a decorative stretch, into a word break.
Does it work for plain English too?
Yes, it is an ordinary slug function for Latin text — separator, lower, maxLength that cuts on a word boundary, stop words and a keep list. The Arabic handling simply does not get in the way.
More from Devix
All resources →Prayer Times
Library
Prayer times and Qibla direction with every calculation method.
VAT Calculator
Library
VAT and GST maths for the UAE, Saudi Arabia and Pakistan.
E-Invoice QR
Library
ZATCA e-invoice QR codes and UAE FTA invoice helpers.