Arabic and English AI Agent Evaluation in the UAE
Arabic and English AI agent evaluation tests whether a system completes the intended business task, uses evidence and tools correctly, respects authority and communicates suitably across the language varieties its users actually employ. Translation of an English test set is only one part of the evidence.
Bilingual deployment creates more than a translation question
An agent can sound fluent in Arabic and still retrieve the wrong policy, misread a customer request, pass unsuitable arguments to a tool or change the meaning of an explanation. It can also complete a task in English while failing when the same customer alternates between Arabic and English.
Evaluation should begin with the business workflow and the people who use or are affected by it. The question is whether the complete system performs acceptably in the language conditions that occur in use, not whether a model can produce grammatical Arabic in a general conversation.
Define the language conditions from the operating reality
Arabic is not one uniform test condition. A UAE system may encounter Modern Standard Arabic, Emirati or other Gulf usage, Arabic speakers from other regions, formal and informal registers, English-Arabic code-switching and variable spelling. Voice channels add accent, audio quality, numbers, names and speech-recognition errors. Latin-script Arabic may matter in some digital channels.
The evaluation plan should select conditions from the intended users, channels and consequences. It does not need every dialect for every use. It does need evidence for the populations and situations in which the organisation intends to rely on the system, and a stated boundary for conditions it does not support.
Use paired and naturally authored cases
Paired English and Arabic cases can test whether equivalent facts and instructions lead to equivalent decisions, tool actions and outcomes. The pair should be reviewed for meaning, ambiguity, tone and policy terms rather than produced through unexamined machine translation.
Not every useful Arabic case should begin in English. Naturally authored cases expose idiom, indirect requests, dialect, cultural context, spelling variation and pragmatic meaning that translation may remove. Native speakers with suitable domain knowledge should author or review material cases and define what an acceptable response means in that workflow.
Test the complete agent pathway
An apparent language failure can begin in several components. Speech may be transcribed incorrectly. A query may be normalised into a different intent. Retrieval may find an English document but miss the applicable Arabic source. A tool may receive a mistranslated category or number. The final response may be correct in content but unsuitable in register, explanation or escalation.
Marketways records the trajectory from input through retrieval, reasoning steps available for inspection, tool calls, state, output and handoff. Component tests then locate whether the weakness lies in recognition, retrieval, policy interpretation, action or communication. This makes remediation more precise than changing the model whenever an end-to-end case fails.
Evaluate task, evidence, language and control separately
The scorecard should distinguish business task completion, factual and policy correctness, evidence support, intent understanding, tool selection and arguments, permission, escalation, language meaning, register and clarity. Safety and consumer-impact criteria should reflect the use.
A single average can conceal a serious difference between language conditions or error types. Results should be reported by task, consequence, language variety and relevant user group with the number of trials and uncertainty visible. A system should not compensate for a material Arabic failure by performing well on a larger set of easy English cases.
Calibrate human and automated graders
Deterministic checks can verify structured outputs, citations, tool parameters and prohibited actions. Model-based graders can help screen larger volumes where their rubric and limitations are tested. Human reviewers remain necessary for material questions of meaning, dialect, cultural pragmatics, domain accuracy and the quality of explanation.
Reviewers should use observable criteria and examples, then grade a shared sample to examine agreement. Differences can reveal an ambiguous rubric, a genuinely context-dependent answer or inconsistent severity. Automated judges should be compared with qualified human judgement in each language condition rather than assumed to transfer from English.
Test switching, ambiguity and recovery
Real interactions do not stay inside a clean benchmark. A user may begin in English, provide an Arabic document, switch dialect, use an English product term and ask for the answer in Arabic. Names, dates, currencies and account references may appear in both scripts.
The test set should include realistic switching, incomplete requests, spelling variation, ambiguous pronouns, corrections, interruptions and requests outside the supported boundary. The agent should clarify, preserve the meaning of material facts, avoid inventing a translation and hand the case to a suitable person when confidence or authority is insufficient.
Connect language quality to customer outcome
A polite and fluent response is not enough if the customer reaches the wrong route, receives an unsupported answer or cannot challenge a consequential outcome. The evaluation should measure completion, correction, repeat contact, escalation, abandonment and complaint patterns as well as linguistic quality.
For high-consequence uses, the organisation should test whether disclosures and explanations are understandable, accurate and equivalent in practical effect. In the financial sector, the CBUAE's 2026 Guidance discusses plain-language disclosures in Arabic and English and measures to check understandability. Institutions should interpret applicable requirements with their qualified regulatory and legal advisers.
Compare systems under the same operating conditions
Vendor demonstrations often use different prompts, settings, retrieval sources and examples. A fair comparison uses the same held-out cases, tools, knowledge boundary, permitted retries and scoring rules. It records model and system versions because performance can change after an update.
The comparison should include a simpler baseline and the existing human or digital route. A model with stronger general Arabic may still be less suitable if it cannot use the organisation's evidence, follow its policy, call tools reliably or support the required deployment and monitoring controls.
Set release thresholds and monitor production
Release criteria should reflect the consequence of each error, not only a global pass rate. A material prohibited action, unsupported eligibility answer or failure to escalate may be a stop condition even when most conversations succeed. Lower-consequence wording differences may be monitored and improved after release.
Production monitoring should sample both languages and the relevant varieties, with additional review after model, prompt, retrieval, policy or tool changes. Complaints, corrections, overrides, unresolved handoffs and unexpected code-switching can become new test cases. Language coverage should follow actual use rather than remain fixed to the launch dataset.
What Marketways delivers
Marketways produces the language and population plan, bilingual and naturally authored test cases, workflow traces, component and end-to-end tests, grader rubrics and calibration, repeated-trial results, disparity and error analysis, release thresholds, remediation priorities and production-monitoring design.
AI Evaluation and Assurance is the canonical service for the engagement. How to Build an AI Agent Evaluation Test Set explains general test-set construction, while AI Agent Evaluation Metrics, Scorecards and Release Thresholds covers the broader measurement design. The bilingual plan adapts those methods to the language, culture and operating conditions of the UAE use.
What to bring to the first discussion
Useful starting material includes the intended workflows and users, channel volumes by language, examples of real requests with suitable protection, supported dialect and language policies, Arabic and English source material, prompts and tools, current model and system versions, human-review procedures, known language failures, complaints, release criteria and people able to provide native language and domain judgement.
