How Unicode Messages work

Helpful insights to explain how Unicode characters work

What is Unicode Standard?

The Unicode Standard is a universal character‑encoding framework created to ensure consistency across digital systems. It supports an extensive range of letters, symbols, and scripts that extend beyond the basic GSM character set. While the GSM alphabet includes Latin characters, numbers, and a limited set of symbols, many writing systems—such as Chinese and Thai—as well as technical symbols, pictographs, and emojis fall under Unicode.

What is a Unicode Message?

A Unicode message is any message encoded using the Unicode Standard. If a message contains even one Unicode character, it must be encoded in this format to ensure those characters display correctly.

How many characters can a Unicode message contain?

A Unicode message can include up to 70 characters per SMS before it is divided into multiple segments.

When a Unicode message exceeds 70 characters and is segmented, each part is limited to 67 characters, because three characters’ worth of space is reserved for the data required to reassemble the message in the correct sequence.

How a Unicode character is represented

Unicode Symbol

Name

‘ ’

Apostrophe

“ ”

Double quotation marks

`

Grave accent

-

Hyphen/dash

&

Ampersand

ß

Eszett (German)

Pilcrow (English document formating and footnotes)

©

Copyright

Ω

Omega (Greek)

÷

Division

Infinity

Ñ

Tilde (Latin, Spanish)

'

Curly Apostrophe (Can be replaced with ')

 

Notes relating to Unicode characters

• Unicode behavior can vary depending on the carrier’s support.

 • If you are sending messages through our HTTP API, Unicode content must be triple‑encoded to ensure it is interpreted correctly. 

• When composing messages in CXT or through integrations (APIs), avoid copying text from MS Word or similar programs. Type the message directly into MXT or your software to prevent unintentionally inserting Unicode characters. 

• When sending Unicode characters via SMPP, ensure your DCS value is set to 8 so the Unicode data is encoded properly.

 

 

 

We use cookies to personalize your experience. By continuing to visit this website you agree to our use of cookies

More