Humanitarian artificial intelligence begins before the model: lessons from KoboToolbox for responsible digital commons
Publication information:
Abstract
Who governs the artificial intelligence deployed in the name of populations in crisis – and who can say no to it? The co-founders of KoboToolbox draw lessons from their own tool for a humanitarian AI governed, and not endured, by those it is supposed to serve.
Artificial intelligence (AI) is no longer just a possibility in the humanitarian sector. It is already present in project proposals, narrative reports, translations, interview summaries, tracking spreadsheets, and coordination notes. It is used to formulate a forward-looking development plan for an organization, to summarize an evaluation, translate a testimony, or analyze open-ended responses.
Full text
AI arrives through utility
AI is spreading because it responds to a real pressure: to do more with less. Needs are increasing, funding is shrinking, reporting requirements are piling up, and teams are becoming exhausted. In this context, it appears as a practical, immediate, almost obvious solution. AI is becoming the technology of humanitarian austerity: a promise of efficiency offered to a sector forced to respond to more suffering with fewer resources [1].
It is therefore no longer its use that is in question. It is already here. The global study conducted in 2025 by Data Friendly Space and the Humanitarian Leadership Academy with 2,539 respondents in 144 countries and territories indicated that 93% of respondents had already used AI tools, while only 22% reported that their organization had a formal AI policy. A follow-up survey conducted in January 2026 confirmed this trend: 95% of respondents were using AI tools, but less than a quarter worked in an organization with a formal policy [2]..
The choice is therefore not between AI and its rejection, but between allowing the most opaque systems to become the default infrastructure for humanitarian action, or building safer, more open and better governed alternatives.
This article does not address all the effects of AI on humanitarian action. It does not cover disinformation, cyberattacks, AI as a civilian infrastructure exposed to conflict, or the full range of geopolitical transformations linked to these technologies. Its focus is narrower: the conditions under which AI enters humanitarian practices, particularly through the collection, analysis, representation of needs, consultation, and decision-making. The central question is less about the performance of the models than about the rights, power, and responsibility in the concrete uses of these technologies.
The lesson from KoboToolbox: infrastructure matters as much as the tool.
The KoboToolbox experience (see box) sheds light on this debate, although the analogy has its limits. KoboToolbox is not artificial intelligence. Its story shows how a humanitarian technology can become essential when it addresses a clear operational problem, while also reminding us that a responsible tool is not defined solely by its practical value. Digital data collection has reduced certain costs, limited data entry errors, improved data quality, and made collection accessible to teams that lacked the resources to develop their own systems. But a free and open-source tool does not automatically become trustworthy. It must operate under challenging conditions, be maintained, allow for data control, and be integrated into training, support, documentation, and governance practices. This lesson applies to AI: a humanitarian technology only becomes responsible through the conditions of its production, access, control, maintenance, and use.
KoboToolbox Kobo is a global, US-based, non-profit organization that develops and maintains KoboToolbox, an open-source platform for collecting, managing, and analyzing data. Designed for challenging environments, including humanitarian crises, development, public health, and human rights contexts, KoboToolbox enables organizations to replace paper forms with simple, reliable, and offline-useful digital tools. Its mission is to make access to quality data faster, more affordable, and more equitable, so that decisions are better informed by the realities of the populations concerned. Used by tens of thousands of organizations worldwide, Kobo functions as a public-interest digital infrastructure for more effective, responsible, and accountable action. |
KoboToolbox and Humanitarian AI
KoboToolbox is a good illustration of, and a test of, this transition, as we are progressively integrating AI into it. This takes very concrete forms: facilitating and accelerating interview transcription, translating responses, processing qualitative data, helping to build forms, and developing intelligent interviewers capable of prompting, asking for clarification, detecting incomplete answers, or suggesting follow-up questions—particularly useful in certain contexts, such as an Ebola outbreak. Some of these features, including transcription, translation, and qualitative analysis of audio responses, are already available.
These uses touch upon the core of humanitarian work, at least with regard to knowledge production. In an investigation, the quality of a response depends on language, context, trust, and the interviewer's ability to recognize nuance. AI that supports transcription, translation, or interview tracking can reduce real workloads and improve certain aspects of data quality. But it can also subtly shift the locus of meaning-making.
While AI helps build forms, translate, transcribe, code responses, or conduct interviews, it also influences what is requested, retained, grouped, or made visible. The challenge, therefore, is to integrate AI without losing the strengths of tools like KoboToolbox: user control, data protection, visible boundaries, human review, documentation, and adaptability to real-world situations.
The risk is that the production of humanitarian knowledge will shift to systems that field teams, local partners, and affected populations cannot inspect, influence, or challenge. AI can support tools, data collection, and analysis. It must not become a silent authority on what deserves to be asked, understood, or retained.
The risk of AI by default
The danger, therefore, is not that humanitarian organizations use AI; it arises when they use it through infrastructures they do not control. In practice, the "default" AI will often be the fastest, the cheapest, the best integrated with existing tools, or simply the only one available. It will also very often be a commercial tool hosted in private infrastructures, governed by terms of service that few organizations actually read, let alone negotiate.
“Digital tools are never neutral instruments: they reshape power relations, institutional practices, and possible forms of action.”
It would be unfair, however, to conclude that the teams are acting irresponsibly. Many use AI with caution and good intentions. But asking people under pressure to reject useful tools without a safe alternative is not a governance policy. It's a shifting of responsibility. If the sector doesn't establish its own conditions of use, others will: suppliers, funders, budgetary constraints, and informal practices.
This risk is not unique to AI. Critical work on humanitarian technologies has already shown that digital tools are never neutral instruments: they reshape power relations, institutional practices and possible forms of action [3]But AI makes this issue more acute, because it does not simply collect or organize information.
When the consultation becomes simulation
AI doesn't just accelerate data collection; it can also interfere with interpretation, representation, and decision-making. Even worse, AI can produce the appearance of a response without anyone being consulted. The "next generation" KoboToolbox isn't intended to replace consultation with generated responses, simulated communities, or automated decisions. Rather, the goal is to leverage AI for tasks it excels at: helping to refine questions, facilitating data collection, accelerating analysis, and organizing responses. AI thus serves to enhance consultation processes, not replace them.
Indeed, humanitarian data is never a purely technical matter. An investigation, an interview, a focus group, a grievance mechanism, or a community consultation are encounters shaped in turn, or simultaneously, by trust, fear, language, power, and institutional expectations [4]Before an AI model can predict needs, classify situations or recommend actions, people's experience has already been transformed into data: a survey response, a GPS coordinate, a photograph, a testimony, a vulnerability category, a need indicator, a complaint or an interview note.
AI then amplifies this transformation. It allows for the much faster processing of large volumes of data, linking them to other sources and extracting categories, profiles, or priorities that were not necessarily anticipated at the time of collection. Data collected to understand a crisis or improve a program can be reused, summarized, combined, modeled, or simulated in unforeseen contexts. The risk, therefore, is not only data extraction. It is intelligence extraction: transforming the experience of populations into models useful to the humanitarian system without giving them power over what these models make visible, invisible, or prioritized. Research on the protection of humanitarian data and metadata already shows that the digital traces produced in humanitarian action can expose those involved to secondary uses, surveillance, or risks that they can neither anticipate nor control [ 5 ].For example, an AI system could analyze protection complaints, survey responses, and location data to assign a risk level to certain households. A misclassification could then influence access to assistance, a follow-up visit, or a protection measure, without the individuals concerned knowing how this classification was produced, nor being able to correct or contest it.
This is where community consultation becomes even more important. Using large language models to simulate consultation with affected communities may seem efficient: why organize a lengthy and costly consultation if a model can produce responses supposedly reflecting a group's likely reactions? But this temptation must be called what it is: a substitution. A model trained to mimic affected populations doesn't give them a voice, while simultaneously providing institutions with a less expensive substitute for listening.
Even when it is only a matter of testing a service or a digital tool, specialists recommend limiting the use of AI-generated profiles, intended to simulate real users, to the formulation of hypotheses, without using them to replace exchanges with real people or to base final decisions.
The value of consultation lies not only in the information it produces, but also in the process itself: the encounter, the listening, the possibility of disagreement, the trust, the recognition of the other as an interlocutor and not merely a source of data. Replacing this relationship with a simulation would be tantamount to confusing informational traces with political presence.
However, affected populations not only have an interest in being better represented, but also the right to be heard, to refuse certain forms of representation, and to challenge the uses made of data or narratives produced from their experience. This requirement aligns with the principles of a human rights-based approach to data, humanitarian commitments to accountability to affected populations, and research on data justice, which emphasizes how individuals are made visible, represented, and treated through the data produced about them.
Judgment cannot be automated.
The use of AI ultimately raises a decision-making question. It is often presented as a tool to aid analysis, prioritization, or planning. This is often true. But in organizations under pressure, this assistance can quickly become delegation. A recommendation generated by a system, especially when integrated into a dashboard or aligned with a donor's expectations, can acquire disproportionate authority.
In humanitarian aid, difficult decisions are not simply a matter of optimization. Who do we help first? Where and with whom do we negotiate access? How do we balance speed, impartiality, safety, dignity, and protection? These questions involve data, but they cannot be solved by data alone. They involve values, responsibilities, risks accepted or rejected, and sometimes the courage to not follow the seemingly most efficient solution.
Part of humanitarian judgment lies precisely in what machines lack: hesitation, doubt, empathy, sometimes even a form of cautious irrationality. These human delays are not always flaws. They can be what prevents a quick decision from becoming irreversible.
AI can identify patterns, generate scenarios, summarize information, flag inconsistencies, and formulate hypotheses. It can help us think. But it cannot assume the moral responsibility for a decision, bear the burden of arbitration, or be accountable to a community. This is why the idea of "keeping humans in the loop" is insufficient if it is not accompanied by a genuine capacity for intervention, challenge, and institutional accountability. Research on the human oversight of automated systems shows that this oversight can become merely symbolic when it fails to consider the automation of practices, deference to systems, or the practical inability of individuals to challenge an algorithmic recommendation [6]..
For responsible humanitarian digital commons
The answer cannot be inaction: the sector cannot prevent humanitarian workers from using tools that meet real needs. The challenge is to facilitate the least dangerous uses and make the riskiest uses more complex.
This implies a shift from a logic of principles to a logic of technological usage frameworks. Responsible humanitarian AI cannot be limited to a charter, training, or a list of prohibitions. It must be embedded in a practical framework that links together the type of use, the level of risk, the data involved, the people concerned, institutional responsibilities, recourse mechanisms, and the technical conditions of deployment. Recent work on humanitarian AI emphasizes precisely this shift from a general ethic to operational mechanisms: risk classification, validation steps, conditions for refusal, post-deployment monitoring, and accountability to the people concerned [7]..
Such a framework must begin by distinguishing between uses. Not all uses of AI carry the same level of risk. There is a significant difference between summarizing a public document and using AI to process sensitive testimony, classify protection complaints, direct assistance, prioritize households, or simulate community preferences. Some uses can be encouraged with simple precautions. Others require enhanced human validation, risk assessment, public documentation, community consultation, or the right to challenge. Some should simply be excluded.
The framework must also address technical requirements. The humanitarian sector doesn't need the most powerful model. It needs models that are sufficiently useful, streamlined, multilingual, documented, auditable, and adaptable to be governed. It needs environments designed for its practices: data classification, contextual warnings, limits on sensitive information, secure hosting, validated corpora, traceability, and the ability to work locally or offline. Openness can help, but it is not enough.
"A framework for use only makes sense if it transforms principles into effective rights and responsibilities into verifiable obligations."
Finally, a framework for use is only meaningful if it transforms principles into effective rights and responsibilities into verifiable obligations. It must define red lines, purchasing rules, incident procedures, audit mechanisms, risk thresholds, and forms of redress. It must guarantee the possibility of genuine human review, the right to refuse certain uses, the right to challenge a decision, and, in some cases, the necessity of consent or collective governance. These guarantees are not merely ethical preferences; they are part of a broader evolution of the legal and regulatory frameworks relating to automated decision-making, human oversight, and risks to fundamental rights.
Humanitarian AI cannot simply increase the capacities of large international organizations. It must also strengthen those of local partners and affected populations: training, resources, accessible tools, less dominant languages, understandable documentation and effective power to challenge [8].
This is where the idea of humanitarian digital commons takes on its full meaning. A commons is not simply a free or open tool. It is an institutional arrangement around a shared resource: who contributes, who accesses it, who decides, who maintains it, who benefits, and who can challenge it. Applied to AI, this idea compels the sector to ask a political question before a technical one: who governs the intelligence generated from crises, and in whose name? This conception is rooted in the tradition of work on the governance of the commons, which demonstrates that shared resources require rules, institutions, and accountability mechanisms tailored to their specific uses.
Humanitarian AI will not be held accountable solely for its usefulness, nor for a few principles added after the fact. It will only become accountable if the sector can shape the conditions under which it is designed, deployed, controlled, and challenged. The goal is not to go faster at any cost, nor to measure, model, or simulate ever more affected populations. It is to build technologies that strengthen the capacity of these populations to be heard, to understand what is being done in their name, to refuse certain uses, and to demand accountability. The question, therefore, is who is served by AI, who governs it, who can challenge it, and who can say no to it.