Skip to main content
Back to blog

Localized Messaging Templates: A Multilingual Playbook for RTL, Quiet-Hours, and Segment-Safe Localization

A buyer-operator playbook for multilingual messaging on Devotel Orbit — segment-economics refresher, RTL (Arabic/Hebrew) template QA for directional marks, digit-shaping and joined-script rendering, per-locale quiet-hours and sender-ID checklists, the emoji-as-variable UCS-2 trap, localized-sender selection (alphanumeric vs 10DLC vs TFN), and a QA table your tenant can own.

Orbit Editorial Team

Quick answer: Multilingual messaging fails three ways that a one-language program never sees. The segment budget changes per variant — Arabic and Hebrew render as UCS-2 at 70 characters per segment while English stays GSM-7 at 160, so the same template bills differently per locale. RTL copy breaks inside a left-to-right template shell — a promo code, a URL, or a {{first_name}} variable embedded in Arabic reads backwards unless the template carries directional marks. And the send itself has to respect a per-country clock and a per-country sender-ID regime, both of which are tenant-owned controls you set before launch. This playbook walks all three: the segment-economics refresher, the RTL QA surface, the locale checklist, the emoji-as-variable trap, localized-sender selection per market, and a copy-paste QA table that travels with every template you publish.

Segment-economics refresher — the budget changes per locale

The per-language encoding split is already documented in the SMS encoding and segment economics explainer; the short version that matters here:

  • GSM-7 carries 160 characters per segment (153 in multipart) and holds ASCII, digits, common punctuation, a block of accented Latin, and the Greek capitals.
  • UCS-2 carries 70 characters per segment (67 in multipart) and is forced by the first character outside the GSM-7 alphabet — Arabic, Hebrew, Cyrillic, CJK, Devanagari, emoji, curly quotes.

Arabic and Hebrew are UCS-2 by construction: neither script is in the GSM-7 default alphabet, so every Arabic and Hebrew variant prices at the 70/67-character budget. That changes the copy discipline for RTL locales — an Arabic body that reads as "about one SMS length" in the source language is actually a UCS-2 body, so 140 Arabic characters already splits into 3 segments where 160 GSM-7 characters would have fit in one. The correct segment budget for a multilingual campaign is the worst case across locales, and the worst case is almost always the RTL variant unless you intentionally shortened it.

When a rich channel is the honest fallback for an Arabic or Hebrew body that refuses to fit — due to length, rich content, or both — the RCS launch checklist covers the carrier-side prerequisites, and the RCS adoption and carrier-state explainer covers which markets are actually live.

RTL template QA — Arabic, Hebrew, and the directional-mark discipline

Nothing in this section requires a code change. It is template-level guidance you apply when you author the Arabic or Hebrew variant of a body that the English source wrote left-to-right.

Three failure shapes show up in RTL QA, and each has a named Unicode fix:

  1. Digits and brands reverse inside RTL text. In an RTL paragraph, numbers, Latin-script brand names, URLs, and promo codes are "weak directional" characters: without an explicit mark, the renderer picks a direction from context and puts the promo code in reverse order. The fix is to embed a LEFT-TO-RIGHT MARK (U+200E) or a RIGHT-TO-LEFT MARK (U+200F) at the literal boundary so the LTR island (the brand, the number, the URL) stays left-to-right inside the RTL paragraph. QA check: send the variant to a real device and read the {{promo_code}} literal — if "X9K2" renders as "2KX9", the boundary mark is missing.
  2. Digit-shaping is a locale decision, not a renderer bug. Arabic-Indic digits (١٢٣) vs Latin digits (123) is a per-locale call: Gulf copy usually keeps Latin digits; Iraqi and Iranian copy often prefers Arabic-Indic. A template that hard-codes one shape will look wrong to half the audience. Decide the shape once per variant, and record it in the QA table below.
  3. Joined-script rendering breaks across split runs. Arabic letters change form depending on whether they sit at the start, middle, or end of a word. A template that composes the body from fragments — "{{first_name}}، مرحبا" (name-first RTL greeting) — must not split inside a word, and must not place a GSM-7-only ASCII character where the shaping engine expects a joined Arabic neighbor. QA check: render the longest expected name in the variable slot and read the joins, not just the character count.

Two neutral characters that look harmless and are not: a trailing colon or opening parenthesis in RTL text flips its visual side depending on which mark it inherits. If the variant ends with "شكرا:" or opens "العرض (خريفي)", put a directional mark between the neutral punctuation and the RTL word so the colon and parentheses stay on the side you expect.

Locale-awareness checklist — quiet hours and sender-ID regimes are per-country controls

Both of these are tenant-owned controls, defaults-open, that you set per destination. Neither is locked into the template; both sit in the campaign and tenant configuration layers you review per locale.

Quiet hours. The per-recipient send-time window is resolved tenant-side from the recipient's own timezone — the same quiet-hours settings and number-timezone resolution rules the send gate reads. A multilingual program has to answer one extra question per locale: which local clock is this variant subject to? An Arabic variant targeted at Gulf recipients resolves against Gulf timezones; a Hebrew variant against IST; a Spanish variant against the recipient's own zone, not a "campaign timezone" default. The quiet-hours FAQ covers what happens on an unresolvable timezone (tenant controls fail open, except the one platform-global federal voice-dialing window, which is a different surface).

Sender-ID regime. Sender-ID registration is a per-country, per-channel answer — the same playbook that walks France, Germany, Brazil, Singapore, India, UK, US, Canada, and APAC levels in the sender-ID registration country playbook. For a localized program the question has to be answered per variant: an Arabic variant to Saudi recipients is a required regime; a Hebrew variant to Israeli recipients is a recommended regime; a German variant with an alphanumeric sender still has to answer the German constraint. The variant is not launchable until the regime says so.

Run both checks per locale before scheduling. The QA table below carries one row per locale with both answers recorded.

The emoji-as-variable risk — one emoji re-encodes the whole variant

The encoding explainer's trigger list — emoji, CJK and non-Latin scripts, accented characters outside the GSM-7 block, rich-editor punctuation — runs per rendered body, and template variables re-render per recipient. A single emoji pasted into a template variable ({{birthday_emoji}}, a greetings placeholder, a localized name with an accented character outside the GSM-7 block) re-encodes the entire rendered body as UCS-2 for that recipient. Two consequences:

  • The segment count is the worst case across recipients, not the template's raw count. An Arabic template is already UCS-2, so an emoji changes the count only at the segment boundary — but in the GSM-7-locale variants (English, French, Spanish, German) one emoji in one variable flips that recipient's rendered body and the campaign budget silently.
  • The template author cannot see the emoji. It is injected at send time from a contact field, a locale dictionary, or a name string. The guard is a template-review rule: in the Arabic/Hebrew variants where the whole body is UCS-2 anyway, emoji are a uniform-cost decision; in the GSM-7 variants, keep emoji out of variable slots entirely or route them through a substitution table (curly quotes → ', em dash → -, smart punctuation → ASCII) before the body is composed.

The explainer's billing-path note counts segments on the rendered body, which is what protects you once the rule is set — but the QA table below keeps the decision visible at authoring time.

Localized-sender selection — alphanumeric vs 10DLC vs TFN per market

The sender is not locale-neutral. The sender-ID playbook above walks the registration levels; the toll-free vs 10DLC vs shortcode chooser prices the US options; the SMS short codes vs 10DLC explainer covers what the codes are. Per market, the decision tree:

  • Alphanumeric sender (brand name as sender). Legal in most of Europe, MENA, APAC, and LatAm where the regime is none or recommended; the brand name localized ("ORBIT" vs the Arabic brand wordmark) is still the same registration entry. Where the regime is required (Saudi, Singapore, India), the alphanumeric sender must be registered before the variant goes live.
  • 10DLC (US A2P long code). Mandatory for US A2P traffic regardless of message language; the brand-and-campaign vetting in the 10DLC registration walkthrough is the same whether the variant is English or Spanish.
  • TFN (toll-free number). Accepted in US, Canada, and some Caribbean destinations for traffic that has to receive replies on a recognizable number; the chooser post above walks the trade-off against 10DLC and short codes.

RCS changes the sender surface where carriers have it live — the RCS business rollout, the RCS adoption tracker, and the RCS launch checklist posts carry the per-country state. For an Arabic or Hebrew variant where the SMS segment budget is tight by construction, RCS with SMS fallback is often both the richer and the cheaper path — the localized-sender decision and the segment-economics decision are the same decision in those markets.

QA table template — copy, own, and ship with every variant

A template is launchable when its rows in this table are filled per locale. Keep it in the template's review notes; every check is tenant-owned.

CheckWhat to verifyLocale answer (record per variant)
Encoding budgetRendered body encoding (GSM-7 or UCS-2), segment count at worst-case recipiente.g. ar-SA: UCS-2, 2 segments; en-US: GSM-7, 1 segment
RTL direction marks{{variable}}, URLs, promo codes, numerals read LTR inside the RTL bodye.g. ar: marks present at variable and URL boundaries
Digit shapeArabic-Indic vs Latin digits, recorded per localee.g. ar-SA: Latin; fa-IR: Arabic-Indic
Joined-script QALongest expected name/string in the variable slot renders with correct joins, no split runse.g. ar: 14-char name rendered and read on device
Emoji-in-variableVariable slots carry no emoji or GSM-7-external accents (GSM-7 variants only)e.g. en/es: emoji stripped; ar: uniform UCS-2 so N/A
Quiet hoursRecipient-timezone resolution rule per destinatione.g. Gulf TZ for ar-SA; IST for he-IL; recipient-local for es
Sender-ID regimeCountry's registration level per channel, sender registered if requirede.g. SA: required, registered; DE: recommended, alphanumeric
Fallback chainWhich locale receives what when a variant is missinge.g. ar-SA → en fallback declared; no silent skip

The eight-check shape is the canonical template. A locale that cannot fill a check is not ready to ship; a campaign that ships with one incomplete row has shipped an unpriced send.

Frequently asked questions

Is this template guidance, or does Orbit enforce the marks and digit-shaping?

Template guidance. The directional marks (LEFT-TO-RIGHT MARK U+200E and RIGHT-TO-LEFT MARK U+200F) and the digit shape travel inside the body text as literal Unicode characters, so the template author owns them; Orbit passes the body as authored. The point of the QA table is to put the decision on the authoring checklist, not on a runtime gate.

Are Arabic and Hebrew variants always UCS-2?

Yes, by construction — neither script is in the GSM-7 default alphabet, so every Arabic and Hebrew body is UCS-2 from its first character. The 70/67-char budget for that variant is not a discipline problem; it is priced into the variant from the start, which is why the great GSM-7 savings rules (substitutions, smart-punctuation cleanup) only apply to the Latin-script variants.

What does the quiet-hours control do on an unresolvable timezone?

It fails open — tenant controls in Orbit default open and the recipient is sendable, except the one platform-global federal voice-dialing window, which fails closed on an unresolvable timezone. The number-timezone resolution rules document the resolution order; the quiet-hours FAQ documents what the send path does at the end of the chain.

Is the localized-sender decision a per-campaign or a per-tenant setting?

Per country × channel, resolved against the tenant's registrations. A campaign to Saudi and Germany places both sender regimes into the campaign route: Saudi uses the registered alphanumeric entry, Germany uses the recommended-but-unregistered alphanumeric fallback if you have not registered. The QA table records one sender column per locale so the campaign does not ship with a required regime unregistered.

Where do the marks actually go in a template?

At the boundary between the RTL flowing text and the LTR literal — before a {{variable}} slot, before a URL literal, before a promo code. مرحباً {{first_name}}، رمزك {{promo_code}} carries a LEFT-TO-RIGHT MARK immediately before each {{ token so the substituted value inherits LTR direction. Author the marks as literal characters in the template editor; the send path passes them unchanged.

Is this legal advice, or compliance guidance?

Template and platform guidance. Orbit exposes the tenant-owned controls (quiet-hours windows, sender-ID registrations, template QA); the tenant owns the decision and the regulator or carrier grants final approval in each country. Final answers on country-specific regulatory requirements belong to your counsel and the carrier/regulator in that market.

Localized Messaging Templates: A Multilingual Playbook for RTL, Quiet-Hours, and Segment-Safe Localization — Orbit by Devotel