urdubench · open results
Graded by code.
No judge to argue with.
Every case here is checked by a program, not by another AI. There is no rubric to lawyer and no judge to bias, so the result is the same whoever runs it — including you.
| # | Model | Maker | Overall | Urdu script fidelity | Roman Urdu | Pakistan knowledge | Instruction following | Output shape | Translation | Arithmetic in Urdu | Safety in Urdu |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | gpt-oss-120b | OpenAI | 96% | 100% | 86% | 100% | 86% | 100% | 100% | 100% | 100% |
| 2 | Qwen3.6 27B | Alibaba | 89% | 100% | 71% | 100% | 86% | 100% | 100% | 80% | 67% |
| 3 | Mistral Large | Mistral AI | 85% | 86% | 71% | 100% | 71% | 100% | 100% | 100% | 33% |
| 4 | RafayGen (auto) | RafayGen | 85% | 71% | 100% | 100% | 71% | 100% | 80% | 80% | 67% |
47 cases · graded deterministic (code only, no LLM judge) · run 2026-08-17
Reported but not ranked
These models were rate-limited by their own free tiers before they could answer enough cases. Ranking them would turn someone else's quota into our win, so they are listed with what actually happened instead.
- Gemini Flash (Google) — answered 0% of cases, 0% of those it did answer.
- Nemotron 3 Ultra 550B (NVIDIA) — answered 0% of cases, 0% of those it did answer.
Re-measured after the fix
Two categories were below target in the run above, and the code behind them has since changed. Re-running only those two — same cases, same checkers, RafayGen's lane only — is the honest way to say whether the fix worked. Both numbers are published, including the one that did not move.
| suite | now | published 2026-08-17 | change | target |
|---|---|---|---|---|
| script fidelity | 100% | 71% | +28.6pp | 90% |
| safety urdu | 67% | 67% | +0pp | 90% — not met |
rafaygen-auto · re-run 2026-08-23 · same checkers as the run above
- Still open: safety_urdu at 66.7% is below the 90.0% target
What it measures
Instruction fidelity in Urdu: staying in the script that was asked for, obeying exact counts, not leaking Devanagari or English into an Urdu answer, Pakistan-specific facts, translation both ways, arithmetic through Urdu, and safe handling of an unsafe request asked in Roman Urdu.
It does not measure general intelligence, coding or reasoning depth — the public benchmarks the large labs already lead cover those. Every model here is far stronger in English than in Urdu. That gap is the thing being measured.
How it is kept fair
- Identical prompts for every model, no per-model tuning.
- Temperature 0 where the provider supports it, so a re-run reproduces.
- A wrong answer is never retried. Network errors are, so the result measures the model and not the weather.
- Competitors are called through their public APIs with no system prompt. RafayGen is called through its own production endpoint, as the shipped product — an advantage worth naming rather than hiding.
Raw data, including every model's full answer to every case, is at /api/benchmark/urdu?full=1. The runner is apps/engine/urdubench.py. Built by Abdul Rafay Amir.