NaijaBench
How well does each AI actually know Nigeria?
A contamination-free benchmark generated from live Nigerian public data — budgets, the knowledge graph, and this week's news. Every question has a verifiable answer, scored automatically (no LLM judge). Frontier models know Nigeria broadly — but they can't answer fresh, specific civic facts (a state's exact budget, today's news). Only a system with the data can.
| # | Model | Overall | Budget | FAAC | Party | Birthplace | City→State | Zone | News (fresh) |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Our agent (data-backed)1000 Reasonsdata-backed | 61% | 67% | 100% | 33% | 67% | 100% | 40% | 33% |
| 2 | Gemini 2.5 FlashGoogle | 43% | 0% | 0% | 0% | 67% | 100% | 100% | 0% |
| 3 | Llama 3.1 8BMeta | 39% | 0% | 0% | 0% | 33% | 100% | 100% | 0% |
| 4 | Mixtral 8x7BMistral | 39% | 0% | 0% | 0% | 100% | 100% | 60% | 0% |
| 5 | Nemotron Super 49BNVIDIA | 35% | 0% | 0% | 0% | 0% | 100% | 80% | 33% |
Fresh / specific facts (exact budgets, FAAC, this week's news) — no model can memorise these; a data-backed system must retrieve them.