Mahak (مِحَكّ)
Open community benchmark evaluating LLM Arabic fluency on authentic everyday tasks
Arabic NLPLLM BenchmarkAstroCloudflare Edge
The Problem
Why This Was Built
General LLM benchmarks (MMLU, Chatbot Arena) rely heavily on translated English benchmarks that fail to evaluate authentic Arabic syntactic flow, legal terminology, and cultural nuance.
The Solution
The Architectural Answer
A community-annotated benchmark scoring models across authentic Arabic prompts with public Elo ranking matrices and native agent evaluation interfaces.
Capabilities
Key Features & Architecture
Blind side-by-side comparison interface and public Elo ranking leaderboard
Live public matrix ranking 65 frontier and open models across 25 authentic tasks
Go CLI (mahak-bench) for headless prompt runs, recovery manifests, and batch synchronization
Model Context Protocol (MCP) server for automated agent benchmark evaluation
Evaluation domains: formal correspondence, contracts, customer support, literature, and instruction following
Engineering
Technical Specifications
- Primary Languages
- TypeScriptGoAstro
- Frameworks & Tooling
- Cloudflare WorkersD1Tailwind CSS
- Architecture & Execution Model
- Edge-rendered ranking matrix, blind side-by-side voter engine, Go CLI, and MCP server
- Licensing Model
- Open Data / Waqf (waqf)
- Release & Lineage
- Active release cycle since 2026 · Maintained by @JadMadi
Knowledge
Frequently Asked Questions
How does Mahak differ from standard benchmarks like ArabicMMLU?
ArabicMMLU tests translated multiple-choice knowledge. Mahak evaluates authentic generation, dialectal nuances, formal legal drafting, and tone adaptation.
Where can I access the live leaderboards?
The public matrix and side-by-side arena are live at https://mahak.waqf.dev.