Arabic AI voice notes: Gulf dialect, MSA and customers who switch to English mid-sentence
An Arabic AI voice note works when it sounds like the customer's own Arabic, not a news bulletin. That means a Gulf voice for Gulf customers and English words left in English where the customer used them. It also means every name, price and time checked by native listeners before launch. Voice is harder than text because Arabic speech carries the dialect in every vowel, and most public Arabic speech data leans towards Modern Standard Arabic. This guide, from Dubai WhatsApp AI builder Learnmind.ai, sets out the research, the WhatsApp rules and a testing method.
The short answer
Voice exposes the dialect. A text reply can sit in neutral written Arabic. A spoken reply cannot, because every vowel and consonant places the speaker somewhere.
Modern Standard Arabic sounds formal in a voice note. It is the Arabic of formal settings. Answering a relaxed Khaleeji voice note in it sounds like a bulletin, even when every word is correct.
Gulf customers mix English in. In one Emirati podcast dataset, 36% of sentences switched between Arabic and English. A good voice reply keeps the customer's English words.
Names, numbers and gender are where voices fail. Written Arabic usually omits short vowels, so a voice has to guess them. "You" also changes form for a man or a woman.
Test with native listeners from each audience. Machines are poor at guessing dialect from audio. In one study, an off-the-shelf model picked the right dialect 36.44% of the time.
Why Arabic voice is harder than Arabic text
Arabic has a formal standard and many spoken dialects. A 2024 Arabic speech benchmark from ELM Company in Saudi Arabia puts it plainly: Modern Standard Arabic "is commonly used in formal settings". The dialects have "notable differences in both pronunciation and vocabulary". Customers write to a business in a mix of both. In everyday speech, dialect dominates.
In text, a business can hide behind the standard. A reply in clean written Arabic reads as polite and neutral to an Emirati, a Saudi and an Egyptian alike. A voice note has no such hiding place. The moment it says the word for "now", it has chosen a region.
Written Arabic adds a second problem. Diacritics, the marks for short vowels, "are routinely omitted in standard written texts", notes a 2026 paper from the Qatar Computing Research Institute (QCRI). So a voice reading علم must decide whether it means knowledge (ʿilm), a flag (ʿalam) or "he taught" (ʿallama). A person decides from context. A machine can get it wrong in a way every listener hears.
| Issue | In a text reply | In a voice note |
|---|---|---|
| Dialect | Neutral written Arabic is widely acceptable | Every word is pronounced one way or another, so a region is always chosen |
| Short vowels | Left out, and the reader fills them in | The voice must pick them, and a wrong pick changes the word |
| Gender of "you" | Often identical on the page, as in بدك ("you want" in Levantine) | Said differently: biddak to a man, biddik to a woman |
| English words | Can stay in Latin letters inside an Arabic sentence | Must be pronounced with an accent that fits the rest of the sentence |
| Numbers | Digits are read the same in every country | Spoken in full, in dialect, with the right noun form after them |
The technology has a third problem: what it was trained on. Researchers at Mohamed Bin Zayed University of Artificial Intelligence note that large Arabic speech datasets "consist mostly of MSA, Egyptian, and Saudi Arabic". The QCRI paper names earlier Arabic text-to-speech datasets of 4, 7 and 12 hours. The largest publicly available one, it says, holds 83 hours, and 73 of those hours are synthesised.
For Emirati speech the gap is wider. In March 2026, Rania Al-Sabbagh of the University of Sharjah introduced Ramsa, a 41-hour Emirati speech corpus. The paper states that, to the author's knowledge, "no published TTS results currently exist for Emirati speech". Its own zero-shot tests set first baselines that, the author says, "leave substantial room for improvement".
A business that wants an Emirati-sounding voice is ahead of the research. It has to test the result itself.
- Modern Standard Arabic (MSA)
- The shared written and formal standard across the Arab world, heard in news and official speech.
- Khaleeji (Gulf Arabic)
- The group of dialects spoken in the Gulf states. Research datasets often label it "Khaliji" or "Gulf", and split it further by country or region.
- Register
- How formal or casual the language is. The same message can be said in MSA, in a softened dialect or in street dialect.
- Text-to-speech (TTS)
- Software that turns a written reply into spoken audio.
- Word error rate (WER)
- The share of words a speech recogniser gets wrong against a human transcript. It can pass 100% when the system adds many words that were never said.
MSA, Gulf, Egyptian and Levantine: what changes when the AI speaks
The differences between dialects are not a matter of accent alone. The MADAR project, published in 2018, had translators from 25 Arab cities render the same travel sentences into their own city's Arabic. Across every pair of cities, with MSA included, the average overlap in vocabulary was 25.8%. The closest pair, Amman and Jerusalem, shared 54.4%. Of all 25 city dialects, Muscat was the closest to MSA, at 37.5% overlap.
One small word shows the scale of it. Here is how MADAR's lexicon records the word for "very" in some of the cities a Gulf business hears from.
| City in MADAR | Words recorded for "very" | Arabic script |
|---|---|---|
| MSA (the concept key) | jiddan | جدا |
| Doha | kellish, waayid, waajid | كلش، واجد |
| Muscat | kithiir, 3oom, waagid | كثير، عوم، واجد |
| Riyadh | jiddan, kithiir | جدا، كثير |
| Jeddah | jiddan, katiir, marra | جدا، كثير، مرة |
| Cairo | giddan, khaaliṣ, 2awi, kitiir | جدا، خالص، قوي، كثير |
| Beirut and Damascus | ktiir | كثير |
Spellings follow MADAR's phonetic transcription, simplified. Note how one written word, كثير, is said four different ways in this table alone. A voice reading it from text has to pick one.
Inside one country: the Emirati sound shifts
Even "Emirati" is not one voice. The Ramsa corpus records three subdialects: Urban, Bedouin and Mountain, also called Shihhi. Shihhi is spoken mainly in Ras Al Khaimah. Ramsa's interviewees also said the boundaries between subdialects are blurring among younger speakers, many of whom use more than one.
Ramsa transcribes Emirati speech as it is said, not as MSA spells it. Its guidelines list sound changes a generic Arabic voice will not make on its own.
| Sound change | MSA form | Emirati form recorded | Meaning |
|---|---|---|---|
| j becomes y | jadīdah جديدة | yadīdah | new |
| q becomes g | ʿuqb | ʿugub | after |
| k becomes sh, for a woman | ḥaḍratik حضرتك | ḥaẓratish | you (polite, to a woman) |
| Final glottal stop dropped | shayʾ شيء | shi | something |
| Phrases fused | mā shāʾ Allāh | mashāllah | God bless |
| Everyday words | al-ān الآن | ilḥīn or ilḥīna | now |
A voice that says jadīdah and al-ān to an Emirati customer is understood perfectly. It also sounds like a voice from somewhere else. That is the gap between a correct reply and one that feels local.
Saudi Arabia has the same layering. The Saudi Audio Dataset for Arabic, known as SADA, labels Najdi, Hijazi and Khaliji speech separately. Riyadh sits in Najd and Jeddah in the Hijaz, and in MADAR only Jeddah's translators used marra for "very".
When an MSA voice sounds like a news reader
The subtler failure in Arabic voice replies is not a wrong word. It is the wrong register. A customer sends a relaxed voice note in Gulf dialect, and the reply comes back in perfect MSA at broadcast pace. The customer understood it, and heard a machine. Here is the same reply in four registers.
One question, four ways to answer it out loud
Customer (Gulf voice note): هلا، عندكم موعد اليوم العصر؟ Hala, ʿindkum mawʿid al-yōm al-ʿaṣr? "Hi, do you have an appointment this afternoon?"
MSA reply: أهلًا وسهلًا بك. نعم، يتوفر لدينا موعد اليوم في الساعة الرابعة عصرًا. هل تودّ أن أحجزه لك؟ Correct, and formal enough for a bank letter.
Gulf reply: هلا والله! إي عندنا موعد اليوم الساعة أربع العصر. تبي أحجزه لك؟ Hala wallah! Ī, ʿindana mawʿid al-yōm, as-sāʿa arbaʿ al-ʿaṣr. Tabi aḥjizah lak?
Egyptian reply: أهلًا بيك! أيوه، عندنا ميعاد النهارده الساعة أربعة العصر. تحب أحجزهولك؟ Ahlan bīk! Aywa, ʿandina miʿād innahārda…
Levantine reply: أهلين! إي في عنا موعد اليوم الساعة أربعة العصر. بدك احجزلك ياه؟ Ahlēn! Ī, fī ʿanna mawʿid…
To a woman: Gulf تبين tabīn, Egyptian تحبي tiḥibbi, and Levantine biddik, which is spelled exactly like biddak.
Illustrative example. Four registers, written by us to show the contrast.None of the four replies is wrong. The MSA reply fits a formal enquiry. The Gulf reply fits a Gulf customer who opened in dialect. The Egyptian and Levantine replies fit expatriate customers who wrote that way first.
Three signs your Arabic voice sounds like a bulletin.
Full case endings on every word, such as ʿaṣran instead of il-ʿaṣr. Formal verbs where a person would use a short one, such as yatawaffar ("is available") instead of ʿindana ("we have"). A steady, even pace with no warmth at the greeting.
Any one of these is fine in a formal context. All three in a reply to a casual voice note will put the customer off.
The fix is a register decision made before launch, not a better voice. Decide the default register and when it should shift. Then write the scripts in it, because a voice can only read what it is given. Our guide to writing AI voice note scripts that sound like a person covers the writing side in detail.
Customers switch to English mid-sentence, and the voice has to keep up
In the UAE, Arabic and English mix inside the same sentence. The Mixat dataset, built from two public podcasts by Emirati hosts, makes the point with numbers. Of its 5,316 sentences, 1,947 contain code-switching. That is 36%. The authors link this to the UAE's expatriate communities, which they note outnumber Emiratis, to bilingual schooling and to English as a lingua franca.
The researchers then ran three widely used speech models on the data. Overall, they wrote, "none of the models provide satisfactory transcriptions for this dataset". All three scored word error rates above 100% on the Emirati segments. One multilingual model translated the Emirati speech into MSA instead of transcribing it. Even the best, pre-trained on MSA, stayed above 100%, "showing that transfer from MSA to Emirati Arabic is still challenging".
Commercial systems vary just as widely. A May 2026 company preprint, not peer-reviewed, tested five commercial transcription services on Arabic-English code-switched speech. On Egyptian Arabic with English, word error rates ran from 13.1% for the best service to 59.7% for the worst. One service did worse on Saudi speech than on Egyptian. The lesson: test your own customers' speech, not a vendor's average.
How a voice reply should handle a mixed sentence
| Customer said | Reply that jars | Reply that fits |
|---|---|---|
| أبي أسوي booking حق الـ facial يوم الخميس "I want to make a booking for the facial on Thursday" | MSA with every English word translated: تم استلام طلب حجز جلسة العناية بالبشرة. Correct, and nobody talks like that. | أكيد! الـ facial يوم الخميس عندنا الساعة ست المسا، يناسبك؟ The customer's word "facial" stays in English. |
| "Can I pay by card? ولا لازم كاش؟" "Or does it have to be cash?" | A long English answer that ignores the Arabic half | أكيد بالكارد، ما يحتاج كاش. Short, in the language the question ended in. |
Our five rules for mixed language, from practice rather than any standard:
- Mirror the main language of the last messageMostly Arabic gets Arabic. Mostly English gets English.
- Keep the customer's English nounsIf they said "booking" or "test drive", say it back that way. Translating it sounds like a correction.
- Say English words with a natural Gulf accentA switch into a strong foreign accent jars as much as a switch into MSA.
- Keep brand names as the brand says themPut them in a pronunciation list once, so every voice note says them the same way.
- Never switch language at a handoverThe person who takes over continues in the customer's language. Our page on Arabic WhatsApp AI assistants in the UAE covers the text side.
Arabizi: Arabic typed in Latin letters
Some customers type Arabic in Latin letters and digits. This is Arabizi. Kareem Darwish's 2013 QCRI paper explains that digits stand in for Arabic letters with no English equivalent: "2" for the hamza, "3" for ʿayn. His examples use "7" for ḥā. So "salam, 3indkum maw3id el7in?" is Arabic for "hi, do you have an appointment now?"
Arabizi has no fixed spelling. Darwish lists five common spellings of the word for "liberty": ta7rir, t7rir, tahrir, ta7reer and tahreer. Some Arabizi words are also English words: in "Ana 3awez aroo7 men America", "men" means "from". His system spotted Arabizi among English with 98.5% accuracy, and converted it to Arabic script with 88.7% accuracy.
A voice note sidesteps the script question. A customer who types Arabizi speaks Arabic but may prefer not to read Arabic script. A spoken reply needs no script at all. Any text that follows stays in the script the customer used.
Understanding the Arabic voice notes customers send
Before an AI can reply out loud, it has to understand the customer's voice note. This is where the dialect gap is easiest to measure.
The Open Universal Arabic ASR Leaderboard, from researchers at ELM Company in Saudi Arabia, tested open speech models on the SADA test set. All models, the authors found, "achieve their best results on MSA, but exhibit a significant decline when applied to dialects such as Egyptian and Khaliji". Here is the best model by dialect.
| Speech in the SADA test set | Best open model's word error rate | Roughly, words wrong in a 20-word voice note |
|---|---|---|
| MSA | 19.23% | 4 |
| Najdi | 36.34% | 7 |
| Hijazi | 36.96% | 7 |
| Egyptian | 40.97% | 8 |
| Khaliji | 48.23% | 10 |
The right-hand column is our arithmetic, not the paper's. The models were tested zero-shot, without extra training on this data, and tuning on dialect speech can improve them. The pattern is the point: Khaliji was the hardest of the five.
Background noise makes it worse. On the same test set, the best model's error rate rose from 37.83% in clean audio to 49.21% in noisy audio. Customers record voice notes in cars, malls and kitchens.
The Casablanca dataset, built from TV series in eight dialects, includes Emirati. On its Emirati test speech, a model fine-tuned on MSA scored a 74.24% word error rate. One fine-tuned on Egyptian scored 67.45%, though neither was tuned for Emirati.
Word error rate overstates the damage to meaning, since a variant spelling of a correct word counts as an error. Even so, these error rates put bookings at risk. Three habits keep that risk small.
- Confirm key details back in writing. A date, a time, a car model or a price heard in a voice note is repeated in text, so the customer can correct it.
- Ask, do not guess. If a word that matters is unclear, the reply asks about that word only. "Did you mean this Thursday or next?" is better than a confident wrong booking.
- Hand over when the stakes are high. A complaint, a medical question or a large order heard in a noisy voice note goes to a person with the audio attached.
We go deeper on this side of the problem in how an AI understands the voice notes your customers send.
Names, numbers and prices: the words a voice gets wrong
Customers forgive a slightly odd phrase. They do not forgive their own name said wrongly, or a price that sounds like a different price. These are the items we would test first, and what to listen for.
| Item | What goes wrong | What to test |
|---|---|---|
| The customer's name | Stored in English letters in the CRM: Aisha, Aysha or Ayesha. The voice reads the spelling, not the name. | The 50 most common first names in your contacts, read in an Arabic sentence |
| Staff names | "Dr Leila" said with English vowels in the middle of Arabic | Every name the agent may mention, written into a pronunciation list |
| Prices | AED 250 read digit by digit, or with the wrong noun form after the number | Every price on the menu, said in full in dialect |
| Numbers 21 to 99 | Arabic says the units first: 25 is "five and twenty", خمسة وعشرين | Prices and quantities with two-digit endings |
| Times | "16:30" read as a figure, or with the wrong part of the day | Times written as people say them: "half past four in the afternoon" |
| Months | Levantine speakers may say تشرين الأول for October where Gulf speakers say أكتوبر | Match the month names to the audience, or say the day and date only |
| Digits in two scripts | A message may contain ٢٥٠ or 250. The speech leaderboard converts Eastern Arabic digits to Western ones before scoring. | Both forms, inside the same reply |
Arabic also changes the noun after a number. From three to ten, the noun is plural: دراهم. From eleven upwards, it returns to the singular: درهم. A voice that says "two hundred and fifty dirhams" with the wrong form sounds foreign at once. The simplest fix is in the script: write prices out in words, in the dialect, so the voice has nothing to guess.
A booking confirmation, written for the voice
What the system holds: Name: Aisha. Thursday 09/10. 16:30. Dr Leila. Deposit AED 250.
What a careless script hands the voice: "Aisha, your appointment is 09/10 at 16:30 with Dr Leila, deposit AED 250."
What a careful Gulf script hands the voice: هلا عائشة، موعدك يوم الخميس تسعة أكتوبر، الساعة أربع ونص العصر مع الدكتورة ليلى. والعربون ميتين وخمسين درهم.
Then, in text: Thursday 9 October, 4:30pm, Dr Leila. Deposit AED 250. Location pin.
Illustrative example. The names and figures are invented.The voice carries the warmth. The text carries the facts the customer will need again. In Arabic, where spoken and written numbers differ so much, that split matters even more.
Formality, honorifics and gender: getting the address right
English has one "you". Arabic has a masculine and a feminine form, and dialects mark the difference in sound as well as spelling. Ramsa records an Emirati feminine "you" ending in "-ish" where MSA has "-ik". A voice note that addresses a woman in the masculine is a small error that every listener notices.
The AI needs the customer's gender before it addresses them directly. It often has it, from a name, a past booking or how the customer refers to themselves. When it does not, the script avoids the gendered verb. If in doubt, write the reply around the booking, not the person.
Gulf phrases that warm a voice note, and when they do not fit
| Phrase | Rough meaning | Fits | Avoid |
|---|---|---|---|
| هلا والله hala wallah | A warm hello | A returning customer who opened casually | The first line of a reply to a complaint |
| حياك الله ḥayyāk Allah | Welcome, literally "may God give you life" | A first greeting, especially to a new customer | Repeating it in every voice note of one conversation |
| أبشر abshir | "Consider it done" | Confirming something the business can actually do | Anything that still needs a manager's approval |
| طال عمرك ṭāl ʿumrak | "May your life be long", a respectful address | A senior or long-standing Gulf client | Every sentence, which turns respect into a script |
| يعطيك العافية yaʿṭīk al-ʿāfya | Thanks for the effort | When a customer has sent documents or photos | As a sign-off when nothing was asked of them |
| إن شاء الله in shāʾ Allah | "God willing" | Natural with any future event | As a hedge on a firm, confirmed booking |
These are our script choices, not research findings. Gendered forms differ too: abshir to a man is abshiri to a woman. Have native speakers from your audience approve the final list.
Formality also depends on the business. A luxury car showroom or a private clinic may speak more formally than a salon or a café. The register can sit anywhere from near-MSA to relaxed dialect. What matters is that it is chosen on purpose and held consistently. Our pages on AI voice notes for luxury brands and AI voice notes for clinics show how that choice changes by sector.
Choosing a voice and register per audience: Emirati, Saudi and expat Arab customers
Who your Arabic-speaking customers are decides the voice, and the Gulf states differ sharply. Figures compiled by the Gulf Labour Markets, Migration and Population programme (GLMM) show nationals are just over half the population in Saudi Arabia and Oman. In Kuwait they are about a third. The UAE and Qatar publish no split by nationality.
| Country | Official language | Nationals in the population, mid-2024 | Local Arabic named in the research we read | Other Arabic-language law | Our starting register |
|---|---|---|---|---|---|
| UAE | Arabic (Constitution, Article 7) | Not officially published | Emirati: Urban, Bedouin and Mountain (Shihhi) subdialects (Ramsa) | Check locally | Softened Gulf voice; mirror other dialects in vocabulary, not accent |
| Saudi Arabia | Arabic (Basic Law, Article 1) | 55.6% | Najdi, Hijazi and Khaliji, labelled separately in SADA | Check locally | Najdi-leaning in Riyadh, Hijazi-leaning in Jeddah; test both if you serve both |
| Qatar | Arabic (Constitution, Article 1) | Not officially published | Doha is one of MADAR's 25 city dialects | Law No. 7 of 2019 on protecting the Arabic language | Gulf voice, approved by Qatari listeners |
| Kuwait | Arabic (Constitution, Article 3) | 32.1% (end of 2024) | Check locally | Check locally | Gulf voice, approved by Kuwaiti listeners |
| Bahrain | Arabic (Constitution, Article 2) | 46.6% | Check locally | Check locally | Gulf voice, approved by Bahraini listeners |
| Oman | Arabic (Basic Statute, Article 3, 1996 text as amended in 2011) | 56.8% | Muscat, the closest of MADAR's 25 city dialects to MSA | Check locally | Gulf voice leaning slightly formal; approved by Omani listeners |
Population shares are GLMM's mid-2024 table, compiled from national statistics offices. The Qatar law covers government bodies' audiovisual and written advertising, and Arabic names for companies and trademarks, as described in a Qatari newspaper column. Nothing in that account addresses how a private business's WhatsApp replies must sound. The last column is our recommendation, not a rule.
In the UAE, Arabic-speaking customers are not all Emirati. The Mixat authors note that expatriate communities outnumber Emiratis. A Dubai clinic may hear from an Emirati family and an Egyptian engineer on the same morning.
| Audience | Starting voice | Register | Watch for |
|---|---|---|---|
| Emirati customers | Gulf voice with Emirati vocabulary | Warm and respectful; slightly formal for first contact | Emirati forms such as ilḥīn for "now"; feminine "-ish" endings; honorifics used sparingly |
| Saudi customers | Gulf voice, Najdi or Hijazi leaning by city | Respectful; abshir and ṭāl ʿumrak can fit | Words that differ between Riyadh and Jeddah, such as marra for "very" |
| Egyptian customers in the Gulf | Gulf voice by default, Egyptian only if the business chooses it | Friendly, direct | Egyptian "g" for "j" (giddan); an imitation that misses can sound like a parody |
| Levantine customers | Gulf or softened neutral voice | Friendly, a little more formal than with Gulf customers | Month names; masculine and feminine forms of biddak |
| Formal or official enquiries | Near-MSA voice | Formal, clear, slower pace | Keep it warm at the greeting, or it reads as a bulletin |
Should the AI switch dialect to match each customer? We are cautious. The Casablanca team tested an off-the-shelf dialect identification model on eight dialects, and it was right 36.44% of the time. Imitating a dialect guessed from one voice note risks parody. We prefer one well-chosen house voice, with mirroring kept to vocabulary rather than accent.
How to test Arabic voice notes with native speakers
No benchmark can tell you how your customers will hear your voice notes. The only reliable check is to play the real output to people who sound like your customers. This is the method we recommend. It is ours, not a published standard.
- Name your audiencesList the Arabic-speaking groups who actually message you: Emirati, Saudi, Egyptian, Levantine. Use your own chat history, not guesses.
- Recruit two listeners per audienceStaff, friends of staff or regular customers. They must be native speakers of that dialect, not fluent learners.
- Write a 40-line test scriptGreetings, prices, times, staff names, customer names, area names, one mixed Arabic-English sentence, an apology and a handover line.
- Generate it in every candidate voice and registerTwo voices in two registers gives four versions of each line.
- Play it inside WhatsApp, on a phoneListen through the phone speaker and through earphones. A file that sounds clean on a laptop can sound different after WhatsApp compression.
- Score it blindMix in a few lines recorded by a real staff member as a benchmark. Listeners should not know which is which.
- Fix and retest the failuresAdd mispronounced names to a pronunciation list. Rewrite lines that sound formal. Retest only what failed.
- Keep sampling after launchEach month, have a native speaker listen to a sample of real voice notes the agent sent.
| Question for each listener | Scale | The line fails if |
|---|---|---|
| Did you understand every word the first time? | Yes or no | Any listener says no |
| Where does this speaker sound like they are from? | Free answer | The answer is not the audience you intended |
| Does it sound like a person or a machine? | 1 to 5 | The average is below 4 |
| Is the tone right for a business messaging you? | Too formal, right, too casual | Most listeners pick either end |
| Was any word, name or number said wrongly? | Note the word | Any name or number is flagged |
Two listeners per audience is a small sample, but enough to catch the worst problems. A mispronounced name or price gets flagged by almost everyone. For subtler questions, such as region and tone, a third listener helps.
What WhatsApp allows, and when to stay in text
Meta's Cloud API documentation is specific. A voice message must be an .ogg file with the Opus codec, mono, up to
16 MB, sent with "voice": true. It then shows the business's profile picture with a microphone icon. Sent as
basic audio instead, it shows a download icon and a music icon.
| WhatsApp rule | Source | What it means for Arabic voice notes |
|---|---|---|
| "The play icon will only appear if the file is 512KB or smaller, otherwise it will be replaced with a download icon." | Meta, audio messages | Keep notes short. On our arithmetic, 512 KB holds about two minutes at 32 kbps. |
| Transcripts appear if the customer has set them to Automatic | Meta, audio messages | The customer may read your note rather than hear it. The words must work as text too. |
| Transcripts "are generated on your device" | WhatsApp blog, 21 November 2024 | Transcription happens on the customer's phone, not on WhatsApp's servers |
| Audio is a service message type inside the 24-hour customer service window | Meta, send messages | A voice note can answer a customer who messaged in the last 24 hours |
| Outside the window, only approved templates; template media headers are image, video, GIF or document | Meta, template components | A voice note cannot open a conversation the business starts. Use a template, then speak once the customer replies. |
| From 1 October 2026, service messages are charged, with 1,000 free per business number each month | Meta pricing | Each voice note is a message. October's UAE service rate is USD 0.0157. |
For the full technical walk-through, see sending and receiving voice notes through the WhatsApp Business API. For costs at your own volume, use our WhatsApp pricing calculator for the UAE or read the UAE pricing guide.
When an Arabic voice note helps, and when text is better
WhatsApp said in March 2022 that its users send 7 billion voice messages on an average day. That is a global figure; we found no Gulf-only breakdown. A voice reply is a normal way to talk on WhatsApp. These are our rules for when to use it.
- Voice back when they sent voice. A customer who records a voice note has chosen to talk.
- Voice for warmth, text for facts. A greeting or an apology lands better spoken. A price, an address or a link belongs in text.
- Short notes win. Twenty to forty seconds covers almost any reply, well inside the 512 KB play button.
- Stay in text for sensitive things. Medical detail, legal wording and anything the customer may forward should be written.
Our comparison of voice notes, text and phone calls goes through these choices in more detail. For the wider business case, see how AI voice notes work for business on WhatsApp.
How Learnmind builds Arabic voice notes
Learnmind.ai is based in Dubai and builds bespoke AI systems for customer conversations. They include WhatsApp AI agents, comment and DM automation, an AI voice receptionist and an AI-first CRM. We work in English and Arabic.
| The problem on this page | What a Learnmind WhatsApp agent does |
|---|---|
| Customers write in Arabic, English or both | Arabic is auto-detected, and the agent replies in the customer's own language and the brand's voice |
| Customers send voice notes | It handles the voice notes customers send, and reads photos and answers about what is in them |
| A text reply feels cold at the wrong moment | It replies with a voice note when the moment calls for it: a warm spoken reply, built to sound like a person, not a script |
| Register and formality | Each agent gets a persona with a name, tone, discovery questions, objection handling and escalation rules, written in days 2 to 5 |
| Wrong answers in any language | Answers come from facts the business approves |
| A voice note that needs a person | It hands over with the whole conversation attached, so the customer never repeats themselves |
| Returning customers | It remembers past conversations, and the business sees everything in an operator view |
The build starts with a 15-minute call on day 1, where we map your services, pricing logic, common questions and brand voice. You prepare nothing. The system is built and tested on real questions before it goes live. Most agents are live from about day 14, and a phrase that sounds wrong is fixed the same day. See how a Learnmind agent goes live and what the WhatsApp agent does.
Alcaz Media, a UAE performance marketing agency, runs a Learnmind assistant on its WhatsApp. It answers leads within seconds, around the clock, qualifies them and books strategy sessions. It handles images and voice notes, and sends the founder a summary with the full transcript. The case study is on our results page.
You can try it first on your own number with the free 14-day WhatsApp AI trial. The trial is free; Meta bills its own message fees directly.
Frequently asked questions
Can an AI send voice notes in Gulf Arabic on WhatsApp?
Yes. WhatsApp accepts voice messages as .ogg files with the Opus codec, within 24 hours of the customer's last message.
Whether it sounds Gulf depends on the voice and the script. Test it with Gulf listeners before launch.
Should an Arabic AI voice note use MSA or dialect?
For most Gulf customer conversations, a softened Gulf dialect. MSA suits formal enquiries and official contexts. A relaxed
voice note answered in MSA tends to sound like a bulletin, even when every word is right.
Why do Arabic AI voices sound robotic or formal?
Written Arabic leaves out short vowels, so the voice has to guess them. Most large Arabic speech datasets are MSA,
Egyptian or Saudi. A March 2026 paper found no earlier published text-to-speech results for Emirati speech.
What happens when a customer mixes Arabic and English in one sentence?
It is normal in the UAE: in one Emirati podcast dataset, 36% of sentences did it. A good reply mirrors the main language
and keeps the customer's English words, such as "booking", in English.
Can an AI understand Arabic voice notes from customers?
Yes, with care. On a Saudi test set, the best open model had a 48.23% word error rate on Khaliji speech, against 19.23% on MSA. A good
system confirms dates, times and prices in text, and asks when a key word is unclear.
Does the AI need to know if the customer is a man or a woman?
Often, yes. In Arabic, "you" and many verbs change form, and some dialects mark it only in sound. If the gender is unknown,
the script is phrased around the booking rather than the person.
How long should an Arabic AI voice note be?
We aim for 20 to 40 seconds. WhatsApp shows a play button only for voice messages of 512 KB or smaller. Prices and
addresses go in a text message after the voice note.
How do I test an Arabic AI voice before launch?
Play a 40-line script to at least two native listeners per audience, inside WhatsApp on a phone. Ask whether they
understood every word, where the speaker sounds from, and whether any name or number was wrong.
Who builds Arabic WhatsApp AI agents with voice notes in Dubai?
Learnmind.ai, a Dubai company, builds bespoke WhatsApp AI agents that auto-detect Arabic, reply in the customer's language
and handle voice notes both ways. Agents go live from about day 14, and there is a free 14-day trial on your own number.
Is an AI voice assistant allowed on WhatsApp?
On our reading, yes, for a business answering its own customers. Since 15 January 2026, Meta's terms restrict AI providers
from offering general-purpose assistants on the platform. An agent serving one business's own customers is a different use,
and that is how Learnmind agents are built.
Want to hear your business answer in Gulf Arabic?
Our WhatsApp AI agents detect Arabic, reply in the customer's own language and send a warm voice note when the moment calls for it.
How we checked this, and what we could not settle
Checked: every research figure comes from the paper it is attributed to, read in full text. WhatsApp format rules come from Meta's developer documentation, read on 2 October 2026. Official languages come from the constitutional texts on Constitute Project. Population shares come from GLMM's mid-2024 table, with sources last accessed 28 March 2026.
Not settled: the research measures speech recognition far more than synthesis. For Emirati speech, it has only first, zero-shot text-to-speech baselines. The commercial benchmark is a company preprint. The Qatar law is described from a newspaper column, not the official text. We could not confirm which languages WhatsApp's voice transcripts support.
Our team checked the October 2026 pricing on 27 September 2026. We could not re-confirm it on Meta's page on the day of writing.
Sources
- ResearchAl Ali and Aldarmaki, Mixat: A Data Set of Bilingual Emirati-English Speech, Mohamed Bin Zayed University of Artificial Intelligence, SIGUL 2024
- ResearchTalafha et al., Casablanca: Data and Models for Multidialectal Arabic Speech Recognition, October 2024
- ResearchWang, Alhmoud and Alqurishi, Open Universal Arabic ASR Leaderboard, ELM Company, December 2024
- ResearchAl-Sabbagh, Ramsa: A Large Sociolinguistically Rich Emirati Arabic Speech Corpus for ASR and TTS, University of Sharjah, March 2026
- ResearchMusleh, Zhang and Darwish, More Data, Fewer Diacritics: Scaling Arabic TTS, QCRI, March 2026
- ResearchBouamor et al., The MADAR Arabic Dialect Corpus and Lexicon, LREC 2018
- ResearchDarwish, Arabizi Detection and Conversion to Arabic, QCRI, June 2013
- ResearchAbdoli et al., Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German, May 2026, company preprint, not peer-reviewed
- ResearchGLMM, GCC: Total population and percentage of nationals and non-nationals (national statistics, mid-2024), Gulf Research Center, sources last accessed 28 March 2026
- PlatformMeta for Developers, Audio messages
- PlatformMeta for Developers, Send messages
- PlatformMeta for Developers, Template components, header formats
- PlatformMeta for Developers, Pricing on the WhatsApp Business Platform, service message charges from 1 October 2026
- PlatformMeta for Developers, AI Providers on the WhatsApp Business Platform, effective 15 January 2026
- PlatformWhatsApp blog, Making voice messages better, 30 March 2022, the 7 billion figure
- PlatformWhatsApp blog, Voice message transcripts, 21 November 2024
- LawConstitution of the United Arab Emirates, Article 7, via Constitute Project
- LawBasic Law of Governance of Saudi Arabia, Article 1, via Constitute Project
- LawConstitution of Qatar, Article 1, via Constitute Project
- LawConstitution of Kuwait, Article 3, via Constitute Project
- LawConstitution of Bahrain, Article 2, via Constitute Project
- LawBasic Statute of Oman (1996, as amended in 2011), Article 3, via Constitute Project
- LawThe Peninsula, New law protecting Arabic language, opinion column by Dr Ahmad Abdulmalik, 27 January 2019, on Qatar's Law No. 7 of 2019
- LearnmindThe Learnmind WhatsApp agent, Arabic auto-detection, voice notes and handover
- LearnmindHow a Learnmind agent goes live, the 14-day build and same-day updates
- LearnmindLearnmind results, the Alcaz Media case study
- LearnmindFree 14-day WhatsApp AI trial
Written by Edmund Gay, Learnmind.ai, Dubai. Figures, platform rules and legal texts are as published in the sources above; the scripts, testing method and recommendations are our own.