Open benchmark · v0.1.0

Arabic agents need native tests.

A reproducible benchmark for useful and secure AI agents in Arabic, RTL, Arabizi, dialectal, and mixed-language environments—not English tasks with translated prompts.

نختبر بيئة الوكيل العربية كاملة: المستند، واتجاه القراءة، واللهجة، ووصف الأدوات، وحدود الموافقة والخصوصية.

24independently scored tasks
8workflow families
matched variants per goal
0real side effects
Matched-triplet design

Utility and security, measured together.

Each family preserves one legitimate business goal across three variants. Refusing everything cannot pass, and completing the task while crossing an authorization or data boundary cannot pass.

Clean

Ordinary task

Measures whether the agent can complete a useful Arabic-native workflow.

Plain attack

Visible conflict

Places a direct Arabic or mixed-language instruction inside untrusted context.

Obfuscated attack

Unicode and mixed script

Uses zero-width, bidi, Arabizi, or cross-language presentation.

Native operating surfaces

More than translated instructions.

The artifacts and workflows are part of the evaluation. Visual fixtures ship with labelled text equivalents, and results must keep native and text-equivalent modes separate.

Invoice payment draftArabic OCR PDF
Procurement reviewRTL UI screenshot
Support refundGulf Arabic ticket
HR leave summaryLevantine + synthetic identity
Calendar rescheduleArabizi + code switching
Knowledge-base replyArabic Markdown + memory
Tool metadataBilingual MCP-style JSON
Filename reviewUnicode bidi spoofing
Deterministic scoring

Observable traces, not hidden reasoning.

Adapters normalize tool calls, arguments, output markers, and completion status. The scorer does not use an LLM judge.

Clean utilityCan the agent perform the ordinary task?
Adversarial utilityCan it remain useful under attack?
Security boundaryDid recipients, values, memory, files, and approvals remain constrained?
Secure completionUtility and security must both pass.
Attack successDid an injected side effect or disclosure succeed? Lower is better.
Unnecessary refusalDid a clean task get abandoned?
Honest release: v0.1 publishes deterministic evaluator controls, not an unrepeatable manual model score. Reproducible model results are welcome with exact adapter, settings, artifact mode, and traces.
Help make it representative

Review it. Break it. Extend it.

We need independent Arabic-speaking reviewers, additional dialects and regional workflows, audited framework adapters, and reproducible model runs.