Arabic text that renders wrong in React Native: mixed direction, digits, and dates
Three bugs from Arabic screens:
Product page متوافقة مع موديلات 2026-2008 (should read 2008-2026)
Verify phone typed ١٢٣٤٥٦, six empty boxes
Checkout التوصيل يوم الجمعة، ١ ربيع الأول (a Hijri month)
None of them is a typo. Each one is a default doing exactly what it was designed to do.
Never trust a default with Arabic text. Set direction per paragraph, fold digits to ASCII before you validate, and choose the calendar explicitly.
This is Day 16 of Ship Native and the second of two Arabic and RTL lessons. Day 15 fixed the layout around the text: the language restart, start and end styles, and arrows. This one fixes the text itself. The code comes from three of my apps: a social feed with racing data, an onboarding flow with a phone code, and a store checkout.
I reviewed the examples against the source projects. I did not run them on a device; the runtime table near the end says what was observed and where.
Bidi in plain words
Unicode gives every character a direction type. Arabic and Latin letters are strong: each one carries its own direction. Digits are weak. Spaces and most punctuation are neutral, so they take their direction from the characters around them.
The Unicode bidirectional algorithm works one paragraph at a time. It finds the first strong character in the paragraph, and that character sets the base direction for the whole paragraph. A paragraph that starts with an Arabic word is right to left even when most of it is English.
React Native gives a single Text one textAlign and one base direction. Put three paragraphs (Arabic, English, Arabic) in one Text inside an Arabic layout and the English paragraph sticks to the right edge. An illustration of that block, with sample data:
سماعة لاسلكية بعزل للضوضاء.
.Free returns within 14 days
متوافقة مع موديلات 2026-2008
Switch the app to English and the Arabic paragraphs break the same way, mirrored.
Detect the direction of each paragraph
The fix starts with a helper that finds the first strong character:
const RTL_REGEX = /[֑-߿יִ-﷽ﹰ-ﻼ]/;
const STRONG_CHAR_REGEX = /[A-Za-z֑-߿יִ-﷽ﹰ-ﻼ]/;
export type TextDirection = 'ltr' | 'rtl';
export const detectDirection = (text: string): TextDirection => {
const firstStrongChar = text.trim().match(STRONG_CHAR_REGEX)?.[0];
if (!firstStrongChar) {
return 'ltr';
}
return RTL_REGEX.test(firstStrongChar) ? 'rtl' : 'ltr';
};
A paragraph with no strong character, such as a line holding only a phone number, falls back to left to right. That is also what the bidi algorithm does with such a paragraph. Split the text on \n and call the helper once per paragraph.
Isolate each paragraph and align it
In my FormattedText component each paragraph becomes its own nested Text. Its content sits between a directional isolate and a pop character, and textAlign comes from comparing the paragraph's direction with the layout direction:
const LTR_ISOLATE = '';
const RTL_ISOLATE = '';
const POP_DIRECTIONAL_ISOLATE = '';
const dir = detectDirection(paragraph);
const paragraphIsRTL = dir === 'rtl';
const textAlign: 'left' | 'right' =
paragraphIsRTL === I18nManager.isRTL ? 'left' : 'right';
return (
<Text
key={`paragraph-${paragraphIndex}`}
style={[styles.paragraphText, {textAlign, writingDirection: dir}]}>
{paragraphIsRTL ? RTL_ISOLATE : LTR_ISOLATE}
{inlineChildren.length > 0 ? inlineChildren : paragraph}
{POP_DIRECTIONAL_ISOLATE}
{paragraphIndex < paragraphs.length - 1 ? '\n' : null}
</Text>
);
The isolate keeps one paragraph's direction from leaking into the next. The comparison looks backwards at first. When the paragraph matches the layout, 'left' is correct because React Native already flips left to the start edge in an RTL layout (Day 15 covers that). When they differ, 'right' puts the paragraph on its own start side. Mentions, hashtags and links still render inside each paragraph because the parsing runs per paragraph before this step.
Year ranges that read backwards
A row on a racing data screen showed career years as a range. On an Arabic device it read 2026-2008. The two years are weak, and the en dash between them is neutral, so the paragraph's right-to-left direction ordered the two numbers backwards.
The formatter wraps the range in left-to-right marks. A left-to-right mark (U+200E) is invisible and behaves like a strong left-to-right letter, so the range has a strong neighbour on both sides. The marks and the dash are literal characters in the source; they are written as escapes here:
export const yearRange = (
from: number | null | undefined,
to: number | null | undefined,
): string => {
if (!from && !to) {
return DASH;
}
if (!from || !to || from === to) {
return `${from ?? to}`;
}
return `${from}–${to}`;
};
The test keeps the reason next to the assertion:
describe('yearRange', () => {
const LRM = '';
it('wraps the range in LTR marks so RTL does not reverse it', () => {
// Without these the row rendered "2026-2008" on an Arabic device.
expect(yearRange(2008, 2026)).toBe(`${LRM}2008–2026${LRM}`);
});
});
A font for each script
Direction is half of the job. The glyphs also need a font drawn for their script. The Liana app picks the family from the layout direction:
export const fontFamily = (weight: FontWeight = 'regular'): string => {
if (!I18nManager.isRTL) {
return AppFonts.RoundedElegance;
}
const map: Record<FontWeight, string> = {
bold: AppFonts.Cairo.Bold,
regular: AppFonts.Cairo.Regular,
// ...the other Cairo weights
};
return map[weight];
};
That holds while every string on screen is in the layout language. In mixed content an English paragraph inside an Arabic layout still gets Cairo. The extension below is mine for this lesson, not code from Liana: pass the detected paragraph direction into the choice.
const paragraphFont = (dir: TextDirection, weight: FontWeight) =>
dir === 'rtl' ? cairoWeights[weight] : AppFonts.RoundedElegance;
What \d matches
In JavaScript, \d matches the ten ASCII digits and nothing else, with or without the u flag. Arabic keyboards type a different set of characters:
| Digits | Code points | Typed by | \d |
|---|---|---|---|
| 0-9 | U+0030 to U+0039 | Latin keyboards | matches |
٠-٩ |
U+0660 to U+0669 | Arabic keyboards | no match |
۰-۹ |
U+06F0 to U+06F9 | Persian and Urdu keyboards | no match |
A code field that strips non-digits with \D deletes every Arabic-Indic digit. The user types six digits and sees six empty boxes, and a manual test on a Latin keyboard never shows it.
Fold, strip, slice
My OTP input folds both ranges to ASCII before stripping anything:
/**
* Keep digits, drop everything else, never exceed the code length.
*
* It has to cope with more than typing: an SMS autofill can paste the whole
* code at once, and an Arabic keyboard sends Arabic-Indic digits (٠-٩), which
* `\d` does not match, so those are folded to ASCII first.
*/
export const sanitizeCode = (raw: string, length: number): string =>
raw
.replace(/[٠-٩]/g, d => String(d.charCodeAt(0) - 0x0660))
.replace(/[۰-۹]/g, d => String(d.charCodeAt(0) - 0x06f0))
.replace(/\D/g, '')
.slice(0, length);
Each fold subtracts the code point of that range's zero, so ٤ (U+0664) becomes 4. The order matters: strip first and the Arabic digits are gone before they can be folded. The final slice covers SMS autofill, which pastes the whole code at once. Pasting ١٢٣٤٥٦ gives 123456 and fills all six boxes.
Tests for keyboards you don't use
// An Arabic keyboard sends Arabic-Indic digits, and SMS autofill pastes the
// whole code at once. Both silently produce an empty field if the filter is
// wrong, and neither shows up in a quick manual test on a Latin keyboard.
describe('sanitizeCode', () => {
it('keeps ASCII digits', () => {
expect(sanitizeCode('1234', 4)).toBe('1234');
});
it('folds Arabic-Indic digits to ASCII', () => {
expect(sanitizeCode('٤١٨٢', 4)).toBe('4182');
});
it('folds extended Arabic-Indic digits to ASCII', () => {
expect(sanitizeCode('۴۱۸۲', 4)).toBe('4182');
});
it('drops anything that is not a digit', () => {
expect(sanitizeCode('4-1 8a2', 4)).toBe('4182');
});
it('never exceeds the code length', () => {
expect(sanitizeCode('418203', 4)).toBe('4182');
});
});
Display and input go opposite ways
Typed numbers travel to the server, so they fold to ASCII. Some numbers shown on screen go the other way: the same app writes its setup wizard's step counter as ١/٨, as its web app does. That needs a second helper:
const ARABIC_INDIC = '٠١٢٣٤٥٦٧٨٩';
// The inverse of `sanitizeCode`. Display and input want opposite
// things here; do not collapse them into one helper.
export const toArabicDigits = (value: string): string =>
value.replace(/\d/gu, digit => ARABIC_INDIC[Number(digit)]);
The calendar is a locale default
The store checkout promises a delivery day. Formatted with ar-SA, the app's runtime chose the Islamic calendar and quoted a Hijri month. The date was valid. The store promises Gregorian days.
The default is not fixed. On my machine, Node 22 with ICU 78 (CLDR 48) resolves ar-SA to the Gregorian calendar. The answer depends on the engine and its ICU data, which is the reason to pin it in the locale tag:
// `ar-SA` defaults to the Islamic calendar, which quoted a delivery date as
// "٥ ربيع الأول". `-u-ca-gregory` fixes the calendar; the Arabic month and
// weekday names are what the site itself quotes.
const AR_LOCALE = 'ar-EG-u-ca-gregory';
/** "Fri 14 Aug": the store's promised day, `days` from today. */
export const deliveryBy = (lang: string, days: number) => {
const d = new Date();
d.setDate(d.getDate() + days);
return toLatinDigits(
d.toLocaleDateString(lang.startsWith('ar') ? AR_LOCALE : 'en-GB', {
weekday: 'short',
day: 'numeric',
month: 'short',
}),
);
};
The digit pass
The app prints Latin digits everywhere. With the calendar pinned, Hermes on iOS still kept Arabic-Indic digits inside the formatted date, and it ignored both -u-nu-latn and the numberingSystem option. I checked that on the simulator while building the checkout. Transliterating after formatting works on any engine:
const AR_DIGIT_START = 0x0660; // ٠
const EXT_DIGIT_START = 0x06f0; // ۰ (Persian/Urdu forms of the same digits)
const toLatinDigits = (s: string) =>
s.replace(/[٠-٩۰-۹]/g, d => {
const code = d.charCodeAt(0);
const start = code >= EXT_DIGIT_START ? EXT_DIGIT_START : AR_DIGIT_START;
return String(code - start);
});
Runtime results
For 14 August 2026, with weekday, day and short month:
| Runtime | Locale or option | Output | Source |
|---|---|---|---|
| Node 22.23, ICU 78, CLDR 48 | ar-SA |
الجمعة، ١٤ أغسطس (Gregorian) |
run for this article, 2026-10-05 |
| Node 22.23, ICU 78, CLDR 48 | ar-SA-u-ca-islamic-umalqura |
الجمعة، ١ ربيع الأول |
run for this article, 2026-10-05 |
| Node 22.23, ICU 78, CLDR 48 | ar-EG-u-ca-gregory |
الجمعة، ١٤ أغسطس |
run for this article, 2026-10-05 |
| Node 22.23, ICU 78, CLDR 48 | ar-EG-u-ca-gregory-nu-latn |
الجمعة، 14 أغسطس |
run for this article, 2026-10-05 |
| Node 22.23 | deliveryBy digit pass |
الجمعة، 14 أغسطس |
run for this article, 2026-10-05 |
| Hermes on the iOS simulator, React Native 0.86 | ar-SA |
a Hijri month | observed in the app, August 2026 |
| Hermes on the iOS simulator, React Native 0.86 | -u-nu-latn and numberingSystem |
Arabic-Indic digits kept | observed in the app, August 2026 |
I did not re-run the Hermes rows for this article, and I have no Android row. Check your own runtime before relying on any default.
Bilingual test checklist
Run each item in Arabic and in English, on iOS and on Android.
- Direction: switch to Arabic, kill the app, relaunch. The language and the layout direction survive (Day 15).
- Layout: margins, paddings and corners use start and end and mirror with the layout (Day 15).
- Icons: back and forward arrows mirror; play, check and search icons do not (Day 15).
- Mixed text: a description with Arabic, English and Arabic paragraphs gives each paragraph its own alignment, with punctuation at the right end.
- Ranges: year ranges read 2008-2026 in both layouts.
- Fonts: Arabic paragraphs use the Arabic family and Latin paragraphs the Latin family, whatever the layout.
- Keyboard: on an Arabic keyboard, type the code. All boxes fill. Repeat with a Persian keyboard if your audience uses one.
- Autofill: paste a full Arabic-Indic code. All boxes fill and nothing overflows.
- Dates: delivery dates show a Gregorian month with Latin digits, in Arabic and in English.
- Persistence of language: after a cold start, every screen above still renders in the chosen language.
Limitations
- Wrapping paragraphs adds isolate characters and changes the children of the outer
Text. ChecknumberOfLinestruncation and anything that readsonTextLayoutlines after the change. keyboardType="number-pad"does not guarantee ASCII digits on every Android keyboard. Sanitize anyway.Intlsupport and its data vary by Hermes version and platform. The digit pass removes one dependency on them; the calendar tag removes another.- A field that accepts free text, such as a search box, needs a decision about which characters to fold. The OTP rule here is safe because a code only ever holds digits.
Watch the lesson
Full video on YouTube · 35-second Short
The video release is planned for 10 October 2026 at 4:00 PM Cairo time. Until release, the embeds and links may be unavailable.
Series
- Previous: Day 15: Switching to Arabic: Layout, Arrows, and the Restart That Lost the Language
- Next: Day 17: Crash Reporting Was Installed, So Why Was It Silent?
Comments