Ordinary task
Measures whether the agent can complete a useful Arabic-native workflow.
A reproducible benchmark for useful and secure AI agents in Arabic, RTL, Arabizi, dialectal, and mixed-language environments—not English tasks with translated prompts.
نختبر بيئة الوكيل العربية كاملة: المستند، واتجاه القراءة، واللهجة، ووصف الأدوات، وحدود الموافقة والخصوصية.
Each family preserves one legitimate business goal across three variants. Refusing everything cannot pass, and completing the task while crossing an authorization or data boundary cannot pass.
Measures whether the agent can complete a useful Arabic-native workflow.
Places a direct Arabic or mixed-language instruction inside untrusted context.
Uses zero-width, bidi, Arabizi, or cross-language presentation.
The artifacts and workflows are part of the evaluation. Visual fixtures ship with labelled text equivalents, and results must keep native and text-equivalent modes separate.
Adapters normalize tool calls, arguments, output markers, and completion status. The scorer does not use an LLM judge.
We need independent Arabic-speaking reviewers, additional dialects and regional workflows, audited framework adapters, and reproducible model runs.