Delegates attending ‘AI Evaluations and Standards in Practice: From Model to Impact’ at the margins of the inaugural UN Global Dialogue on AI Governance hosted by IDinsight alongside Permanent Missions of Kenya and The Gambia, Kenya's Office of the Special Envoy on Technology, and the UN Foundation.
Governments across the globe are embedding AI in public services: healthcare triage, agricultural advisories, social protection systems, citizen-facing portals, and beyond. Many are simultaneously drafting governance frameworks, principles, and safeguards. But at the AI for Good Global Summit 2026 and the UN Global Dialogue on AI Governance, hosted this year in Geneva, the conversation kept returning to AI evaluations.
Government officials, UN bodies, and heads of state reiterated the lack of a dependable way to evaluate the AI systems in front of them. The UN’s inaugural AI Scientific Panel also stated this gap in their preliminary report: no solid frameworks for multi-agent evaluation, thin methods for risk mapping, almost no longitudinal study, and little guidance on decision-making under uncertainty.
Addressing this gap, IDinsight co-convened a roundtable on the margins of the inaugural UN Global Dialogue on AI Governance along with the Permanent Missions of Kenya and The Gambia, Kenya’s Office of the Special Envoy on Technology, Lawyer’s Hub Kenya and the UN Foundation. Co-moderated by H.E. Ambassador Philip Thigo (Kenya’s Special Envoy on Technology), Dr Claire Melamed (Vice President, AI and Digital Cooperation at UN Foundation) and Linda Bonyo (Lawyer’s Hub Kenya), the session brought together government bodies, implementers, and civil society to discuss the role of AI evaluations in the current development sector landscape.
To anchor the discussion, we presented the ‘Generative AI Evaluation Playbook’, co-developed with The Agency Fund and Center for Global Development, as a case study. The following four reflections stood out from the discussion in the room.
A recurring theme in the room was how abstract AI evaluation tends to be. We heard variations of the same questions from different officials:
“What criteria should we use to judge the systems we procure? How do we evaluate what a tool actually does, rather than evaluating AI in general?”
In our session, we shifted the focus from evaluating AI as a tool in itself to asking, “Does this specific tool, in this specific service, help the people it is meant to serve?” Our research on how AI evaluation works in practice illustrates what this looks like in specific contexts:
The unit of evaluation has to be the specific application doing a specific job for a specific population, not “AI” as a category. This is achieved by testing the actual tool, in its actual context, against outcomes that matter.
AI systems are not static and cannot be validated or fixed only at the initial design and procurement stages. They change over time: models get updated, usage patterns drift, populations shift, and performance that was fine last quarter can degrade in the next one. Our implementer research showed that because AI systems evolve continuously, teams stopped treating them as fixed interventions to be studied over long horizons and instead relied on frequent, directional signals that support ongoing judgment.
For governments, that has a direct consequence. Evaluation cannot be a one-time priority at design and procurement. It has to be a continuous standing capacity to keep asking critical questions on whether a deployed system is still delivering. That means intermediate proxies a ministry can track now — consultation time, task-completion rates, override frequency — rather than waiting years for a definitive impact study that the technology will have outrun.
Implementers already work this way out of necessity; governments should build it in by design.
Governments need a source of judgment they can trust. One that is not the vendor selling the system or the organisation running it. Independent evaluation bodies are emerging to fill that role, providing quality assurance on the AI governments procure and deploy. But independence alone doesn’t guarantee useful evaluation. An independent evaluator that merely reproduces a vendor’s own performance claims is worse than none at all, because it certifies those claims as validated and gives officials unwarranted confidence in the result.
Another risk is an evaluation that optimises for a passing compliance report rather than for realised service delivery. A system can satisfy every documented requirement and still leave patients waiting or applicants wrongly excluded. When the evaluation targets the wrong metric such as procedural conformance rather than delivered outcomes, it does not simply miss the system’s failures; it masks them, suppressing the very signals a government most needs to detect.
So, what keeps evaluation honest? IDinsight’s implementer research points to two things. First, domain experts—clinicians, legal specialists, subject-matter practitioners—were consistently described as indispensable for catching failures in high-stakes settings. Automated dashboards alone don’t surface them. Second, the most informative signals sit at the messy intersection of product behaviour and user response: where people hesitate, make manual corrections, build workarounds, or quietly disengage. Those are exactly the signals a rubber-stamp process is built to ignore. Practical evaluation is designed to surface them, because they are where service delivery actually breaks.
A question came up repeatedly in Geneva through government stakeholders: who actually controls the AI running inside our public services?
Many governments have worked hard to secure data sovereignty, keeping citizen data within national borders and under national law. That effort matters. But it is a mistake to assume that control over the data confers control over the AI. A government can host data domestically and still depend entirely on models it cannot inspect.
AI sovereignty is the capacity to independently answer: is this tool serving our citizens, and what happens when it fails? It is what turns a government from a standards-recipient, certifying compliance against criteria handed down from elsewhere, into a standards-shaper that defines “good” on the terms of its own service-delivery mandate. This is sovereignty that can be built through consistent evaluation practices.
As AI becomes embedded across healthcare, education, agriculture, and public administration, the focus must shift from performance to outcomes: do AI systems improve lives, and are they safe in the conditions where they are actually deployed? We will continue to codify our evolving knowledge of evaluation approaches and tools into our living AI Evaluation Playbook, in partnership with the Centre for Global Development and The Agency Fund, with support from the Gates Foundation.
To learn more about our work in AI evaluations, feel free to reach out to our teammates Poornima Ramesh (poornima.ramesh@idinsight.org) and Charlene Migwe (charlene.migwe@idinsight.org)
—————————————————————————————————————————
The roundtable was co-hosted by the Permanent Missions of the Republic of Kenya and the Republic of The Gambia, and convened by the Office of the Special Envoy on Technology, Republic of Kenya; IDinsight; and the UN Foundation. It was co-organised by the Africa AI Policy Lab, a programme of the Lawyers Hub, and Macmillan Keck, Attorneys & Solicitors. We thank all partners for their collaboration in making this session possible.
7 August 2026
28 July 2026
27 July 2026
21 July 2026
14 July 2026
10 July 2026
9 July 2026
7 July 2026
2 July 2026
11 June 2025
19 September 2025
25 February 2026