Real brief
Every test begins with work a regional builder, team, or business might genuinely need shipped.
المعيار المفتوح للذكاء الاصطناعي في الشرق الأوسط
Most AI leaderboards test for the lab. YarBench tests whether a model can actually code, reason, use tools, and create for the Middle East.
$ yarbench run code/rtl-commerce
→ loading bilingual brief...
→ region: gcc · locale: ar-SA
→ masking model identities...
✓ run record sealed before scoring
_
WHY YARBENCH
The region shouldn’t have to guess which global model understands its work.
Generic benchmarks rarely test the details that break real products here: Arabic mixed with English, right-to-left layouts, Hijri dates, local currencies, unfamiliar addresses, or a brief shaped by the way teams actually work.
YarBench makes those details the test. Not a regional footnote—a first-class standard.
THE INDEX
A single overall rank hides too much. YarBench shows where each model is genuinely useful—and where it falls apart.
Can it ship production-ready software for the region—not just pass a toy coding test?
SAMPLE TASK FAMILIES
SCORE WEIGHT
HOW IT WORKS
A YarBench rank should be challengeable, rerunnable, and useful long after the announcement post disappears.
Every test begins with work a regional builder, team, or business might genuinely need shipped.
Model names and vendor tells are removed. Output order rotates to reduce position and brand bias.
Prompts, harness settings, artifacts, costs, judge cards, and failures stay linked to the result.
Models change. Each result names the exact model, date, harness, region, and benchmark version.
40% objective checks + 35% blind human review + 15% agent reliability + 10% cost & speed
LAUNCH ROSTER
Every family enters unranked. Scores appear after verified public runs—not because a model is popular, new, or good at marketing.
No synthetic launch scores. The first leaderboard will be generated from published run records.
THE NON-NEGOTIABLES
Modern Standard Arabic, Gulf dialects, code-switching, and RTL behavior are scored separately.
A polished demo that breaks on local data does not outrank a plain result that ships.
The same model can perform differently in an IDE, agent, API, or chat. We record the setup.
No score appears without a run record. No leaderboard quietly rewrites its own history.
OPEN BY DEFAULT
The protocol, schemas, task cards, and future run records live in public. Fork it. Challenge it. Add a task from your market. A regional standard only works when the region can shape it.
Open the repository