United Nations

DELLM Methodology and Benchmarking Study

Methodology and benchmarking for evidence-linked language-model workflows.

Project overview

What this project was built to help with

The UN benchmarking study focused on how large language models should be evaluated before deployment in global development contexts. The methodology addressed local-language relevance, data equity, transparency, UN values, Sustainable Development Goal alignment, risk, ethical use, and the practical question of how agencies can judge value and use cases.

Project story

A responsible evaluation frame for development-sector LLMs

UN and development-sector teams needed a clearer methodology for assessing large language models before deploying them in high-stakes contexts. The work focused on more than model performance: it considered local-language relevance, data equity, transparency, alignment with UN values, SDG relevance, risk, and ethical use.

The methodology paper helped clarify how agencies can evaluate LLM value and decide where the technology is appropriate. It also surfaced relevant use cases and gave teams a more disciplined way to discuss responsible AI deployment in development work, especially where evidence quality, local context, and accountability matter.

Example deliverables

Concrete outputs or working assets the project produced or supported.

Deliverable 01Methodology paper for evaluating LLMs in global development
Deliverable 02Benchmarking considerations for value, risk, and ethical use
Deliverable 03Review criteria covering language, equity, transparency, UN values, and SDGs
Deliverable 04Use-case framing for responsible evidence-linked AI workflows
Researchers reviewing methodology and benchmarking materials

Decision problem

UN and development-sector teams needed a clearer methodology for evaluating large language models before using them in global development contexts. The work had to address technical value, ethical risk, local-language relevance, data equity, transparency, SDG alignment, and practical use-case assessment.

What the work involved

1

Developed a methodology paper on how LLMs should be evaluated and deployed in global development settings.

2

Addressed local-language relevance, data equity, model transparency, UN values, Sustainable Development Goal alignment, risks, and ethical use.

3

Clarified how agencies can assess LLM value and identify responsible use cases before relying on evidence-linked AI workflows in high-stakes contexts.

Outcome

A stronger methodology base for responsible evidence-linked AI workflows.

Evidence inputs

What the workflow brings together

Use-case materialsBenchmarking criteriaEvaluation methodsEvidence-linked output examples

Review focus

What stays visible for human review

Methodological rigor
Benchmark validity
Responsible use constraints

Talk with us about a similar workflow.

Bring one decision, report, review process, or evidence problem. We will help identify the smallest useful place to begin.

Start the conversationExplore AidInsight
Cookies

We use cookies to enhance your browsing experience on our website. By continuing to use our site, you consent to our use of cookies. For more information, please review our Privacy Policy.