FOUNDING PROTOCOL · V0.1

BUILT HERE.
TESTED HERE.

المعيار المفتوح للذكاء الاصطناعي في الشرق الأوسط

Most AI leaderboards test for the lab. YarBench tests whether a model can actually code, reason, use tools, and create for the Middle East.

yarbench / run-001

$ yarbench run code/rtl-commerce

→ loading bilingual brief...

→ region: gcc · locale: ar-SA

→ masking model identities...

MODEL AARTIFACT READY
MODEL BARTIFACT READY

✓ run record sealed before scoring

_

MENA-FIRST REAL-WORK TASKS OPEN EVIDENCE AR + EN

WHY YARBENCH

The region shouldn’t have to guess which global model understands its work.

Generic benchmarks rarely test the details that break real products here: Arabic mixed with English, right-to-left layouts, Hijri dates, local currencies, unfamiliar addresses, or a brief shaped by the way teams actually work.

YarBench makes those details the test. Not a regional footnote—a first-class standard.

THE INDEX

One benchmark.
Five kinds of real work.

A single overall rank hides too much. YarBench shows where each model is genuinely useful—and where it falls apart.

TRACK 01 PROTOCOL READY

Code Bench

Can it ship production-ready software for the region—not just pass a toy coding test?

SAMPLE TASK FAMILIES

  • Arabic + English product brief
  • RTL interface repair
  • SAR checkout and tax logic
  • Hijri/Gregorian scheduling

SCORE WEIGHT

Correctness35%
Regional fit25%
Craft15%
Agent reliability15%
Cost + speed10%

HOW IT WORKS

No mystery score.
Follow the evidence.

A YarBench rank should be challengeable, rerunnable, and useful long after the announcement post disappears.

01

Real brief

Every test begins with work a regional builder, team, or business might genuinely need shipped.

02

Blind run

Model names and vendor tells are removed. Output order rotates to reduce position and brand bias.

03

Proof attached

Prompts, harness settings, artifacts, costs, judge cards, and failures stay linked to the result.

04

Versioned score

Models change. Each result names the exact model, date, harness, region, and benchmark version.

YBREGIONAL UTILITY SCORE

40% objective checks + 35% blind human review + 15% agent reliability + 10% cost & speed

LAUNCH ROSTER

The ladder starts
at zero.

Every family enters unranked. Scores appear after verified public runs—not because a model is popular, new, or good at marketing.

MODEL FAMILYTRACKSSTATUSRATING
OGPTOpenAI
AWAITING VERIFIED RUN#01
AClaudeAnthropic
AWAITING VERIFIED RUN#02
GGeminiGoogle
AWAITING VERIFIED RUN#03
XGrokxAI
AWAITING VERIFIED RUN#04
DDeepSeekDeepSeek
AWAITING VERIFIED RUN#05
QQwenAlibaba
AWAITING VERIFIED RUN#06

No synthetic launch scores. The first leaderboard will be generated from published run records.

THE NON-NEGOTIABLES

01

Arabic is not one test

Modern Standard Arabic, Gulf dialects, code-switching, and RTL behavior are scored separately.

02

Useful beats impressive

A polished demo that breaks on local data does not outrank a plain result that ships.

03

The harness matters

The same model can perform differently in an IDE, agent, API, or chat. We record the setup.

04

Receipts before rankings

No score appears without a run record. No leaderboard quietly rewrites its own history.

OPEN BY DEFAULT

Don’t trust the badge.
Inspect the run.

The protocol, schemas, task cards, and future run records live in public. Fork it. Challenge it. Add a task from your market. A regional standard only works when the region can shape it.

Open the repository