WhatsApp-first customer service with AI agents, one memory across WhatsApp, voice and the app
Where the customer writes on WhatsApp first and calls when it matters, the memory has to follow the person across channels, languages and vendors. What the WhatsApp agent needs, how Arabic and English in the same conversation are handled, and what the voice agent in the contact center receives.
Channels7 min read
In a WhatsApp-first operation, the conversation starts in writing and moves to a call when it matters: the customer sends three messages at night, gets an answer from the WhatsApp agent in the morning, and calls the contact center at noon when the answer did not solve it. For the voice agent, or the human agent, to pick up where WhatsApp left off, the memory has to be outside both channels and both vendors, and it has to recognize the same person by the wa_id on one side and the phone number on the other. That is what Niadra does: every message and every call become events of one memory, and each agent reads the context of the person before it answers, in the language the customer used. This post goes through what the WhatsApp agent needs, what happens when the customer mixes Arabic and English in the same conversation, and what the voice agent receives when the phone rings.
Why WhatsApp first changes the memory problem
In markets where WhatsApp is the first channel, like Brazil, India and the Gulf, three things are different from a web chat.
The conversation lasts months. There is no session that ends when the tab closes: the same thread holds the order of March, the complaint of August and the question of today. A memory keyed by session loses the thread; one keyed by the person has to cut that thread into sessions by inactivity, with a deadline per channel, so that "what just happened" is the last exchange and not the whole year. Niadra closes a session by inactivity and extracts the memory of each one when it closes; the whole thread stays searchable.
The identifiers are the channel's. The Cloud API hands over a wa_id (the phone with digits only) and, with WhatsApp usernames, a business-scoped user id that is only unique inside your WhatsApp Business account. The voice platform hands over the number in E.164 and, with STIR/SHAKEN, the carrier's attestation. The CRM has its own contact id. Niadra stores each one as a handle of its own type, with a weight, and merges them only on evidence: a system event, an OTP, your own records. A username is a display attribute and never identifies anyone.
The vendors differ. The WhatsApp agent often comes from a messaging platform, the voice agent from a voice platform, and the agents that work inside the ERP from the company's own team. None of them opens its memory to the others. The memory has to sit above all of them, which is the whole reason Niadra exists.
What the WhatsApp agent needs from the memory
Before it answers, the agent reads the context of the person: the facts, the open items and the promises with the date and the channel they came from, what other agents did inside the company, and the highlights of the history, in chat format. The WhatsApp Cloud API adapter opens the conversation with the wa_id as subject, checks Meta's signature and records each inbound message by its wamid, so a webhook Meta redelivers days later is written once. The reply is recorded by the id the Cloud API returns. Through Twilio, the Twilio adapter reads the WaId on the Messaging webhook the same way.
A message proves that it came from that number, and the recorded turn raises the session to V1. Data your policy reserves for a higher level (a balance, a document, a card) stays withheld until the proof arrives. The proof can come in the same channel: an OTP sent and confirmed on WhatsApp raises the conversation to V3, and the next read brings what that level releases. The WhatsApp agents guide goes through the identifiers, the sessions, the OTP and the media.
Media travels by reference. A voice note or a document arrives as an id, a MIME type and a SHA-256; your code downloads the bytes from the Graph API and hands them over bound to the customer, so erasing the person erases the recording too. The text of a PDF or an image is read inside the region only, and only for the types your space lists.
Arabic and English in the same conversation
A Gulf customer writes in Arabic, switches to English for the product name and the order number, and goes back to Arabic. What the memory does with that, mechanism by mechanism:
- Extraction, context and search work in any language. The model that extracts the memory reads the conversation as it was written. The summary of each episode, the solution, the descriptions of the open items and the subject label come out in the language the customer wrote, even when it is not the space's language; the fixed labels of the context ("Facts", "Open items") follow the language of your space, Portuguese or English.
- Identifiers and numbers are found in any script. The order number, the phone, the amount and the date in a turn are read by exact match, whatever surrounds them. "Where is order 4471?" in Arabic finds the conversation about order 4471.
- The word rules are tuned for three languages. The lexicon that reads a turn to pick the slots of the context, with stemming and synonym groups, covers Portuguese, English and Spanish, and so do the relative time words ("last week", "semana passada"). A turn written in Arabic gets the exact matches and, when your space switches it on, the semantic channel; it does not get the stemmed matches or the time windows. That is the honest limit today.
- Masking runs before any model. E-mail, phone, card, IBAN, postal code and address are masked in any country before the text reaches the extraction model; the ID formats masked by rule are those of one country, and an Emirates ID number is not among them. The extraction never stores a document, card or bank account number in any case.
The result for the voice agent is a context in your space's language with the customer's own words preserved where it matters: "the customer reported the router restarts every night" stays in Arabic if that is how it was said, next to the order number and the date.
When the customer calls the contact center
The phone rings with the carrier's attestation and the number in E.164. Niadra recognizes the same person as the wa_id of the morning, because the two handles were merged on evidence, and the first read starts while the phone rings, within 1.5 s. The voice agent gets the voice view: who is calling, what is open, what the WhatsApp agent answered and what the internal agent did, in a context of 88 tokens at the median, before it says the first word. A message written on WhatsApp while the call is already on reaches the next turn of the call as a delta: the measured freshness from a write on one channel to a read on another is 62.7 ms at the median.
When the agent transfers to a person, the human agent on the desk reads the brief view, a short handover in the size of one utterance, through the tool the contact center already uses: by API, webhook or MCP. Niadra has no agent desk and never will; the screen belongs to your contact center platform. The voice agents guide and the post on the latency budget of memory for voice have the numbers of each deadline.
What this post does not claim
Niadra runs in a single region, stated in the contract, and the data stays there; the only thing that leaves it is already masked text, sent to the AI providers on the subprocessor list. Where that region is for your company is a contract question, and this post makes no claim of local presence anywhere. The word rules in Arabic are a limit, said above, not a feature.
How Niadra solves it
Niadra is the omnichannel memory layer: the WhatsApp messages, the calls, the app sessions and the events of your systems become one memory of each person, recognized by the identifiers each channel has, and any agent, from any vendor, reads the context of its task before it answers, in the language the customer used, at the level the conversation has proven. The adapters for WhatsApp Cloud API, Twilio, LiveKit, Vapi and the others wire it into the platforms you already run. The post on keeping context from a voice call to WhatsApp covers the way back.
Frequently asked questions
Does it work through a WhatsApp provider rather than the Cloud API?
Through Twilio, with the Twilio adapter, which reads the WaId on the Messaging webhook. With other providers, the SDK is used directly: the conversation opened with the wa_id as subject, and the inbound and outbound messages recorded as turns with the provider's message id as idempotency key.
What if the customer switches between Arabic and English mid-conversation?
The memory is extracted from the conversation as it was written, and the summaries keep the customer's language. Identifiers and numbers are found in any script. The stemmed word matches and the relative time words of the context slots are tuned for Portuguese, English and Spanish; an Arabic turn relies on exact matches and, when switched on, the semantic channel.
Does the voice agent in the contact center see the WhatsApp conversation?
It does, before it says the first word, when the number that calls and the wa_id that wrote were merged on evidence: a system event, an OTP or your own records. The WhatsApp exchange appears under "what just happened" in the voice view, and the human agent reads the same memory as a short handover.