The AI tools we recommend for nonprofits – and why
Photo credit: Dali Images
This blog was originally published by The Agency Fund on Jun 25, 2026 — see here.
As social impact organizations increasingly turn to AI, the available tools are evolving faster than most organizations can navigate. For nonprofits operating with small technical teams, it can be especially difficult to compare and select the best AI tools for their needs.
After years of embedding engineering teams with high-impact organizations—and co-designing AI systems alongside them—The Agency Fund has developed a perspective on this challenge. Drawing on partnerships with more than 30 nonprofits across Asia, Africa, and Latin America, we recently conducted a mapping exercise evaluating more than 50 tools.
This blog shares the results: a set of default tool recommendations for NGOs with small technical teams building out a modern technology stack. Our recommendations have been developed in collaboration with data.org, Fast Forward, and IDinsight.
No single tool will work for every organization and context, but the goal is to provide vetted defaults that help social impact organizations get started. While primarily intended for organizations with understaffed engineering teams, the mapping may also help more mature organizations building new components of their technical stack. If you already have a strong engineering team and a modern stack, there are better resources for you. This is for the rest.
For small tech teams, the time and cost of evaluating tools is often a major bottleneck. Teams can spend months comparing platforms and testing integrations before anything reaches users. Good defaults help organizations launch quickly, then adapt or replace tools over time if needed.
Our mapping rated tools using three criteria, taking into account the operational realities facing small engineering organizations in the social sector. Tools were rated on a three-tier scale within each criterion:
The mapping also aligns with the four-level framework for AI evaluation, particularly those first three levels: evaluating AI model performance (L1), product performance (L2), and changes in users’ thoughts and behavior (L3). Broader impact evaluations (L4) fall outside the scope of this mapping, but remain essential: the goal is not simply to deploy effective AI tools, but to understand whether they are generating measurable improvements in users’ lives.
Most organizations will not need to build their entire technology stack at once. The goal of this mapping is not to encourage adoption of every tool below, but to help organizations identify investments that can unlock progress. Where are your operational bottlenecks? For some organizations, it might be data infrastructure or frontline data collection; for others, mobile chat delivery or experimentation capacity.
We evaluated tools across three layers of the stack. Below, we provide default recommendations for each layer. Where relevant, we also highlight strong alternatives, emerging tools, and practical heuristics for choosing among tools based on organizational context, use case, and operational constraints. All these tools are either open source, or offer very generous free tiers and if you don’t want to self-host, many tools below have cloud offerings with nonprofit discounts.
What does it include? Technologies that power AI products, including large language models (LLMs), speech systems, monitoring, evaluation, and retrieval.
This layer sits between your app and the foundation model providers, letting you switch between model systems (e.g., Gemini, OpenAI, Claude) without rewriting code.
Tools that record what your AI is saying and doing, so you can spot when it starts going wrong. These tools can also be used for ad hoc evaluation of model performance in target languages or domains.
Tools that help you find the best speech-to-text and text-to-speech models for your target users across different languages and domains.
Frameworks for building AI agents that take multi-step actions (e.g., searching, fetching information, calling other tools) instead of just answering one question at a time.
These tools store the documents your AI needs to reference, so it can pull the right context when answering questions about your data.
What does it include? Tools that help organizations manage, monitor, and improve AI systems through data pipelines, dashboards, workflow orchestration, analysis, and experimentation.
This is the plumbing that moves data from wherever you collect it (e.g., survey forms, chatbots, apps) into one place you can analyze.
Visual reports that show how your programs are performing across usage, outcomes, or anything else you’re tracking.
Automation for the tasks your team currently does manually: “when X happens, do Y.”
A searchable index of all the data your organization has, so people can find what exists without asking around.
One-off explorations of your data – not regular reports, but specific questions (e.g. “how did this group respond last month?”) that you only want to answer once.
Tools for running A/B tests on your programs: randomly varying what users see, then measuring what works better.
What does it include? Tools that shape how users see and experience your AI products. Most nonprofit products reach users through one of three delivery models: frontline worker services, mobile chatbots, or custom applications built by small teams or technical founders.
For when field staff carry phones or tablets, log data with beneficiaries, and sync it back to your systems – often in low-connectivity and form-heavy contexts. These tools are foundational; most organizations start here.
(Note: The Agency Fund has built lightweight tools to integrate SurveyCTO and CommCare data into organizations’ data systems with minimal effort. Reach out if your data collection layer is in place, but the integration layer is not.)
These tools reach users directly on WhatsApp, Telegram, or by phone with no frontline worker in the loop– for when you’re ready to deliver services at scale, not just collect data.
When you need a web or mobile app of your own – built by your team, designed for your specific workflow.
(Note: These tool recommendations for custom applications aren’t intended for engineering teams of three or more. There are better resources for that audience. But if you’re a solo technical founder or a very small team building a custom application, the stack below is what we’d start with. Together, these tools provide a strong balance between flexibility and ease of use, with documentation and community support that make them practical for small teams.)
Over the course of deploying and evaluating 50+ tools, five broader lessons emerged about how organizations can get the most value from AI systems.
This mapping is a living document. The ecosystem shifts constantly as new tools emerge, older ones are acquired, prices change, and capabilities expand. Our recommendations will continue to evolve alongside it.
We are watching several trends particularly closely: how non-technical users interact with data in the next generation of tooling, how the capacity bar continues to fall for nonprofits building their own products, and how the gap between “tools for big tech” and “tools that fit an NGO” continues to shape what is actually usable in our sector.
As our view sharpens, we will continue updating the map and sharing what we learn.
10 July 2026
9 July 2026
7 July 2026
2 July 2026
29 June 2026
23 June 2026
10 June 2026
7 May 2026
24 May 2014
1 March 2019
7 March 2019
2 April 2019